HomeDatasetsEarthSpeciesProject/NatureLM-audio-training
N

EarthSpeciesProject/NatureLM-audio-training

Audio Classification · EarthSpeciesProject· 15.7K
other 47 GB

Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.

Open in MLForge Sign up free Desktop app
# download instantly
mlforge datasets pull EarthSpeciesProject/NatureLM-audio-training

Dataset details

Task
Audio Classification
Language
en
License
other
Size
47 GB
Creator
EarthSpeciesProject
Downloads
15.7K
Source
huggingface_datasets
Updated
2025-06-03

About EarthSpeciesProject/NatureLM-audio-training

NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained on this dataset may respond with "Common yellowthroat".