ECCV Marine Vision Workshop 2026 · Oral Presentation

WildFin: An In-the-Wild Video Dataset for Fish Behavioral Recognition

Abigail Grassick, Jerome Tze-Hou Hsu, Ethan Lin, Ziang Liu, Max Whitton, Madelyn Hair, Liam Gutierrez, Haozheng Yu, Kristin Branson, Vivek Jayaraman, Michael A. Gil, Andrew M. Hein, Jennifer J. Sun

Cornell University

Abstract

Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments.

WildFin characterizes these failures through a dataset collected and annotated by ecologists. It includes two critical deployment paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The benchmark evaluates modern vision foundation models and quantifies the remaining gap between current model capabilities and real-world underwater behavioral analysis.

Dataset Details

Fieldwork
1,350 hours
Expert annotation
600 hours
Behavioral video
9 hours
Labels
2M+ frame-by-frame annotations
Video settings
Stationary group footage and diver-follow individual footage
Task
Fish behavioral recognition in complex underwater video

Benchmark

WildFin stress-tests static and spatiotemporal vision architectures in marine environments, emphasizing real deployment conditions such as occlusion, changing visibility, camera motion, and multi-animal scenes.