CVPR · 2023 · Conference paper

SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding

Haoxiang Wang, Pavan Kumar Anasosalu Vasu, Fartash Faghri, Raviteja Vemulapalli, Mehrdad Farajtabar, Sachin Mehta, Mohammad Rastegari, Oncel Tuzel, Hadi Pouransari

University of Illinois Urbana-Champaign · Apple (United Kingdom)

Rendered by arXiv from the LaTeX source. If anything looks wrong, switch to the PDF.HTML · PDF · Open in a new tab ↗ · Close
Published in
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
Date
2024-06-17
Citations
180
arXiv
2310.15308
Cite
@inproceedings{wang2023samclip,
  title = {SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding},
  author = {Haoxiang Wang and Pavan Kumar Anasosalu Vasu and Fartash Faghri and Raviteja Vemulapalli and Mehrdad Farajtabar and Sachin Mehta and Mohammad Rastegari and Oncel Tuzel and Hadi Pouransari},
  year = {2023},
  booktitle = {2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)},
  eprint = {2310.15308},
  archivePrefix = {arXiv},
  doi = {10.1109/CVPRW63382.2024.00367},
  url = {https://arxiv.org/abs/2310.15308},
}