ECCV Marine Vision Workshop 2026 · Oral Presentation
WildFin: An In-the-Wild Video Dataset for Fish Behavioral Recognition
Cornell University
Abstract
Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments.
WildFin characterizes these failures through a dataset collected and annotated by ecologists. It includes two critical deployment paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The benchmark evaluates modern vision foundation models and quantifies the remaining gap between current model capabilities and real-world underwater behavioral analysis.
Dataset Details
- Fieldwork
- 1,350 hours
- Expert annotation
- 600 hours
- Behavioral video
- 9 hours
- Labels
- 2M+ frame-by-frame annotations
- Video settings
- Stationary group footage and diver-follow individual footage
- Task
- Fish behavioral recognition in complex underwater video
Benchmark
WildFin stress-tests static and spatiotemporal vision architectures in marine environments, emphasizing real deployment conditions such as occlusion, changing visibility, camera motion, and multi-animal scenes.