Detail publikačního výsledku

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

KUMAR, S.; SEDLÁČEK, Š.; LOKEGAONKAR, V.; LÓPEZ, F.; YU, W.; ANAND, N.; RYU, H.; CHEN, L.; PLIČKA, M.; HLAVÁČEK, M.; ELLINGWOOD, W.; UDUPA, S.; HOU, S.; FERNER, A.; BARAHONA, S.; BOLAÑOS, C.; RAHI, S.; HERRERA-ALARCÓN, L.; DIXIT, S.; PATIL, S.; DESHMUKH, S.; KOROSHINADZE, L.; LIU, Y.; PERERA, L.; ZANOU, E.; STAFYLAKIS, T.; CHUNG, J.; HARWATH, D.; ZHANG, C.; MANOCHA, D.; LOZANO-DIEZ, A.; KESIRAJU, S.; GHOSH, S.; DURAISWAMI, R.

Originální název

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Anglický název

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Druh

Stať ve sborníku v databázi WoS či Scopus

Originální abstrakt

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly “from the wild” rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems’ progression toward audio general intelligence.

Anglický abstrakt

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly “from the wild” rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems’ progression toward audio general intelligence.

Klíčová slova

Artificial intelligence; Audio acoustics; Benchmarking; Intelligent agents; Open systems

Klíčová slova v angličtině

Artificial intelligence; Audio acoustics; Benchmarking; Intelligent agents; Open systems

Autoři

KUMAR, S.; SEDLÁČEK, Š.; LOKEGAONKAR, V.; LÓPEZ, F.; YU, W.; ANAND, N.; RYU, H.; CHEN, L.; PLIČKA, M.; HLAVÁČEK, M.; ELLINGWOOD, W.; UDUPA, S.; HOU, S.; FERNER, A.; BARAHONA, S.; BOLAÑOS, C.; RAHI, S.; HERRERA-ALARCÓN, L.; DIXIT, S.; PATIL, S.; DESHMUKH, S.; KOROSHINADZE, L.; LIU, Y.; PERERA, L.; ZANOU, E.; STAFYLAKIS, T.; CHUNG, J.; HARWATH, D.; ZHANG, C.; MANOCHA, D.; LOZANO-DIEZ, A.; KESIRAJU, S.; GHOSH, S.; DURAISWAMI, R.

Vydáno

20.01.2026

Nakladatel

Association for the Advancement of Artificial Intelligence

Místo

Singapore

ISBN

9781577359067

Kniha

Proceedings of the Aaai Conference on Artificial Intelligence

Svazek

27

Číslo

40

Strany od

22688

Strany do

22697

Strany počet

10

URL

BibTex

@inproceedings{BUT212468,
  author="{} and Šimon {Sedláček} and  {} and  {} and  {} and  {} and  {} and  {} and Maxim {Plička} and  {} and  {} and Sathvik {Udupa} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and  {} and Santosh {Kesiraju} and  {} and  {}",
  title="MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence",
  booktitle="Proceedings of the Aaai Conference on Artificial Intelligence",
  year="2026",
  volume="27",
  number="40",
  pages="22688--22697",
  publisher="Association for the Advancement of Artificial Intelligence",
  address="Singapore",
  doi="10.1609/aaai.v40i27.39430",
  isbn="9781577359067",
  url="https://ojs.aaai.org/index.php/AAAI/article/view/39430"
}

Dokumenty