arXiv · 2025 · Preprint

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception

Miguel Carvalho, Hélder Dias, Bruno Martins

Rendered by arXiv from the LaTeX source. If anything looks wrong, switch to the PDF.HTML · PDF · Open in a new tab ↗ · Close
Date
2025-11-25
Citations
8
arXiv
2511.19820
Cite
@misc{carvalho2025cropvlm,
  title = {CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception},
  author = {Miguel Carvalho and Hélder Dias and Bruno Martins},
  year = {2025},
  eprint = {2511.19820},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2511.19820},
}