ICASSP · 2025 · Conference paper

From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs

Yuhang Jia, Xu Zhang, Yujie Guo, Yang Chen, Shiwan Zhao

Published in
IEEE International Conference on Acoustics, Speech, and Signal Processing
Date
2025-08-03
Citations
0
arXiv
2508.01659
Cite
@inproceedings{jia2025from,
  title = {From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs},
  author = {Yuhang Jia and Xu Zhang and Yujie Guo and Yang Chen and Shiwan Zhao},
  year = {2025},
  booktitle = {IEEE International Conference on Acoustics, Speech, and Signal Processing},
  eprint = {2508.01659},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2508.01659},
}