Skip to content

Latest commit

 

History

History
217 lines (199 loc) · 8.85 KB

File metadata and controls

217 lines (199 loc) · 8.85 KB
layout home
title MARINA Dataset
description A large labeled dataset for underwater acoustic target recognition.
hero
background_image subheadline headline text buttons
/assets/images/coming-soon-background.jpg
Open benchmark · Underwater Acoustics
<strong>UniqueShip:</strong> Large, public underwater acoustic target recognition (UATR) datasets for ships
2,460 hours of ship-radiated noise from 4,218 unique vessels, split by vessel ID so no ship appears in both training and test split. Sourced from the Ocean Networks Canada (ONC) repository.
label url style
Browse releases
#releases
primary
label url style
Read the overview
#overview
outline
overview
subheadline headline paragraphs cta_section stats
Overview
Built to train generalizable UATR models using leakproof splits
UniqueShip pairs hydrophone recordings from seven ONC deployments in the Strait of Georgia (May 2016 – November 2023) with AIS vessel tracking data. Each 5-second sample is labeled with its vessel class and 17 AIS metadata fields. Unlike earlier ONC-based datasets, every split keeps each vessel in a single partition and groups background audio by day, so test accuracy reflects performance on ships the model has never heard.
paragraphs buttons
<strong>The paper</strong> describes the dataset in more detail and includes results with and without data leakage, ablations on vessel diversity vs. audio duration, a metadata analysis, and additional baseline results.
label url target style
Read the paper
_blank
primary
value label
3,437
Hours of ship & background audio
value label
4,218
Unique vessels
value label
11
Vessel classes
value label
2.5M
5-second recordings
difference
subheadline headline items
What makes it different?
Larger, more diverse, and free of data leakage that inflates other benchmarks
title text
Leak-free splits
Vessel audio is grouped by MMSI and background by day instead of random splitting. On previous datasets, random splitting inflated accuracy by 10–48 points.
title text
Largest open ONC dataset
The balanced benchmark subset alone has 4× the audio and 12× the vessels of DeepShip, and 70% more audio than the unbalanced Oceanship dataset.
title text
Rich AIS metadata
17 fields per sample, including MMSI, distance to hydrophone, speed, course, length, beam, draught, and navigation status.
title text
Ready-made splits
Choose anything from a 25-hour quick-start subset to the full 3,437-hour corpus, with five 80/10/10 folds.
title text
Baselines included
MobileNetV3, ViT-B/16, and SwinV2 with STFT and Mel inputs. The best result is 66.5% accuracy (Swin + Mel).
title text
Cleaner background class
8km ship-free radius ensures quieter ambient samples for the background class
releases
subheadline headline text items
Data releases
Current dataset splits
All splits are vessel-disjoint and include per-sample AIS metadata. Samples are 5-second clips at 20 kHz; full-length recordings are available through the codebase. *Request Google Drive permission to access the current splits*.
title image image_alt description size views date buttons
5 Class - Balanced
/assets/images/5shiptypes.png
Aerial view of five vessel classes tracked in open water
Contains the main 5 classes (Tug/Tow, Tanker, Passengership, Cargo) and balances the total audio for each class such that they are equal. Current version = 1.0
89 GB (Unzipped), 59 GB (Zipped)
213h, 3175 vessels, 5 classes
Released September 2026
label url target style
Download spectrograms
_blank
primary
title image image_alt description size views date buttons
12 Class - 5 Hours Each
/assets/images/moreshiptypes.png
Diverse vessels tracked in a busy coastal shipping channel
Contains all ship classes and balances the total audio such that it is 5 hours each class. Current version = 1.0
10 GB (Unzipped), 7 GB (Zipped)
60h, 4218 vessels, 12 classes
Released September 2026
label url target style
Download spectrograms
_blank
primary
leaderboards
subheadline headline status text preview
Benchmark results
Leaderboards
Coming soon
Compare published results across the UniqueShip dataset splits. Rankings, evaluation metrics, and submission guidance will be available following the dataset release.
columns rows
Rank
Submission
Dataset split
Score
Updated
4
inquiries
subheadline headline text benefits form
Inquiries
Need additional information or a different split?
Current dataset splits are available to <a href="#releases" class="text-primary hover:underline">download directly</a> — no request or approval is required. Use this form if you have questions, need additional information, or would like to request a split that is not currently available.
Ask questions about the data or documentation
Request additional or specialized dataset splits
action submit_label submitting_label error_message success fields
Submit inquiry
Submitting…
We could not send your request. Check your connection and try again.
eyebrow headline text
Request received
Thanks for your submission
We have received your request and will be in touch.
type name label autocomplete required width
text
entry.925008751
Full name
name
true
half
type name label autocomplete required width
email
entry.2144715529
Work email
email
true
half
type name label autocomplete required width
text
entry.241770130
Organization
organization
false
full
type name label required width
textarea
entry.1262269016
Inquiry
true
full
citation
subheadline headline text code contact
Citing this dataset
Reference the dataset paper
If you use UniqueShip, please cite the paper below.
@inproceedings{hashemi_2026_uniqueship, author = {Hashemi, Connor and Stout, Trevor and Hoogs, Anthony and Parham, Jason}, title = {UniqueShip: Mitigating Data Leakage in Acoustic Ship Classification Benchmark Datasets}, booktitle = {OCEANS 2026}, year = {2026}, pages = {TODO} }
headline text button_label url
Questions, corrections, or collaboration?
Our team can help with access and research partnerships.
Contact the team