Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,821 to 14,880 of 18,849

  1. Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

    Ke Wang, Junting Pan, Weikang Shi +3

    cs.CVcs.AIcs.CLarXiv:2402.14804v12024
  2. Stereo World Model: Camera-Guided Stereo Video Generation

    Yang-Tian Sun, Zehuan Huang, Yifan Niu +4

    cs.CVarXiv:2603.17375v12026
  3. VideoAtlas: Navigating Long-Form Video in Logarithmic Compute

    Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi +3

    cs.CVcs.AIarXiv:2603.17948v12026
  4. 3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model

    Hyun-kyu Ko, Jihyeon Park, Younghyun Kim +2

    cs.CVarXiv:2603.18524v12026
  5. In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing

    Xiao Fang, Yiming Gong, Stanislav Panev +4

    cs.CVarXiv:2603.19456v12026
  6. LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation

    Junjue Wang, Zhuo Zheng, Ailong Ma +2

    cs.CVarXiv:2110.08733v62021
  7. TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

    Yan Zeng, Haoran Jiang, Kaixin Yao +4

    cs.CVarXiv:2603.17735v12026
  8. Benchmarking Denoising Algorithms with Real Photographs

    Tobias Plötz, Stefan Roth

    cs.CVarXiv:1707.01313v12017
  9. Context-Aware Crowd Counting

    Weizhe Liu, Mathieu Salzmann, Pascal Fua

    cs.CVarXiv:1811.10452v22018
  10. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

    Yixin Liu, Kai Zhang, Yuan Li +9

    cs.CVcs.AIcs.LGarXiv:2402.17177v32024
  11. Robust Global Structure-from-Motion via View Graph Pruning

    Jiamin Xu, Lixing Yao, Weichen Dai +4

    cs.CVarXiv:2608.22054v12026
  12. BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

    Sheng Zhang, Yanbo Xu, Naoto Usuyama +21

    cs.CVcs.CLarXiv:2303.00915v32023
  13. Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders

    Shang-Jui Ray Kuo, Paola Cascante-Bonilla

    cs.CVcs.LGarXiv:2603.19209v12026
  14. Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection

    Youbin Kim, Jinho Park, Hogun Park +1

    cs.CVarXiv:2603.21944v12026
  15. Generalized ODIN: Detecting Out-of-distribution Image without Learning from Out-of-distribution Data

    Yen-Chang Hsu, Yilin Shen, Hongxia Jin +1

    cs.CVcs.LGeess.IVarXiv:2002.11297v22020
  16. Event-Based Motion Estimation via Oriented Distance Fields

    Lei Sun, Yuqin Ma, Weilun Li +5

    cs.CVcs.ROarXiv:2608.24223v12026
  17. VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

    Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas +2

    cs.CVcs.AIcs.LGarXiv:2603.23495v12026
  18. Scaling up GANs for Text-to-Image Synthesis

    Minguk Kang, Jun-Yan Zhu, Richard Zhang +4

    cs.CVcs.GRcs.LGarXiv:2303.05511v22023
  19. Unified Focal loss: Generalising Dice and cross entropy-based losses to handle class imbalanced medical image segmentation

    Michael Yeung, Evis Sala, Carola-Bibiane Schönlieb +1

    eess.IVcs.CVcs.LGarXiv:2102.04525v42021
  20. F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting

    Injae Kim, Chaehyeon Kim, Minseong Bae +2

    cs.CVarXiv:2603.21304v22026
  21. Scaling Out-of-Distribution Detection for Real-World Settings

    Dan Hendrycks, Steven Basart, Mantas Mazeika +5

    cs.CVcs.LGarXiv:1911.11132v42019
  22. Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR

    Sheethal Bhat, Bogdan Georgescu, Awais Mansoor +5

    cs.CVarXiv:2608.24281v12026
  23. Generalized Zero- and Few-Shot Learning via Aligned Variational Autoencoders

    Edgar Schönfeld, Sayna Ebrahimi, Samarth Sinha +2

    cs.CVcs.AIcs.LGarXiv:1812.01784v42018
  24. 4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video

    Jae Won Jang, Yeonjin Chang, Wonsik Shin +2

    cs.CVarXiv:2603.21618v22026
  25. Revisiting Batch Normalization For Practical Domain Adaptation

    Yanghao Li, Naiyan Wang, Jianping Shi +2

    cs.CVcs.LGarXiv:1603.04779v42016
  26. Robotic Pick-and-Place of Novel Objects in Clutter with Multi-Affordance Grasping and Cross-Domain Image Matching

    Andy Zeng, Shuran Song, Kuan-Ting Yu +18

    cs.ROcs.CVarXiv:1710.01330v52017
  27. TrajLoom: Dense Future Trajectory Generation from Video

    Zewei Zhang, Jia Jun Cheng Xian, Kaiwen Liu +4

    cs.CVarXiv:2603.22606v12026
  28. BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

    Risa Shinoda, Kaede Shiohara, Nakamasa Inoue +3

    cs.CVarXiv:2603.23883v12026
  29. RealMaster: Lifting Rendered Scenes into Photorealistic Video

    Dana Cohen-Bar, Ido Sobol, Raphael Bensadoun +5

    cs.CVarXiv:2603.23462v12026
  30. DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models

    Jaewon Min, Jaeeun Lee, Yeji Choi +7

    cs.CVarXiv:2603.23499v12026
  31. Mobile-Former: Bridging MobileNet and Transformer

    Yinpeng Chen, Xiyang Dai, Dongdong Chen +4

    cs.CVcs.LGarXiv:2108.05895v32021
  32. Non-Local Recurrent Network for Image Restoration

    Ding Liu, Bihan Wen, Yuchen Fan +2

    cs.CVarXiv:1806.02919v22018
  33. PixelSmile: Toward Fine-Grained Facial Expression Editing

    Jiabin Hua, Hengyuan Xu, Aojie Li +4

    cs.CVcs.AIarXiv:2603.25728v12026
  34. TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans - A Case Study in Craniosynostosis 3D Photography

    Sanjay Bhandari, Nawazish Khan, Alzbeta Novotna +7

    cs.CVcs.AIarXiv:2608.22131v12026
  35. Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling

    Ruixing Zhang, Hanzhang Jiang, Leilei Sun +3

    cs.CVcs.AIarXiv:2603.26610v12026
  36. One View Is Enough! Monocular Training for In-the-Wild Novel View Generation

    Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez +1

    cs.CVarXiv:2603.23488v22026
  37. Structural Graph Probing of Vision-Language Models

    Haoyu He, Yue Zhuo, Yu Zheng +1

    cs.CVarXiv:2603.27070v12026
  38. SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision

    Avigail Cohen Rimon, Amir Mann, Mirela Ben Chen +1

    cs.CVarXiv:2603.24036v22026
  39. DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery

    Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas +1

    cs.CVcs.CLcs.IRarXiv:2608.21529v12026
  40. Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration

    Danil Tokhchukov, Aysel Mirzoeva, Andrey Kuznetsov +1

    cs.CVarXiv:2603.24800v22026
  41. Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math

    Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6

    cs.AIcs.CLcs.CVarXiv:2603.24961v12026
  42. EfficientFormer: Vision Transformers at MobileNet Speed

    Yanyu Li, Geng Yuan, Yang Wen +5

    cs.CVarXiv:2206.01191v52022
  43. Deep learning with noisy labels: exploring techniques and remedies in medical image analysis

    Davood Karimi, Haoran Dou, Simon K. Warfield +1

    cs.CVcs.LGeess.IVarXiv:1912.02911v42019
  44. MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models

    Bocheng Zou, Mu Cai, Mark Stanley +2

    cs.CVarXiv:2603.25744v22026
  45. DeepStereo: Learning to Predict New Views from the World's Imagery

    John Flynn, Ivan Neulander, James Philbin +1

    cs.CVarXiv:1506.06825v12015
  46. Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification

    Dangwei Li, Xiaotang Chen, Zhang Zhang +1

    cs.CVarXiv:1710.06555v12017
  47. Unified Number-Free Text-to-Motion Generation Via Flow Matching

    Guanhe Huang, Oya Celiktutan

    cs.CVarXiv:2603.27040v12026
  48. CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets

    Robert N. Spaans, Catherine Chia, Tongjie Wang +7

    eess.IVcs.CVarXiv:2608.21497v12026
  49. Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

    Peixi Wu, Ke Mei, Feipeng Ma +15

    cs.CVarXiv:2604.22280v32026
    Summaries:한국어
  50. Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks

    Jie Hu, Li Shen, Samuel Albanie +2

    cs.CVarXiv:1810.12348v32018
  51. VirtualHome: Simulating Household Activities via Programs

    Xavier Puig, Kevin Ra, Marko Boben +4

    cs.CVcs.AIcs.LGarXiv:1806.07011v12018
  52. MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

    Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy +2

    cs.CVcs.AIarXiv:2608.24107v12026
  53. Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis

    Ming Cheng, Hongyu Sun, Zhaolin Chen +3

    cs.CVcs.MMarXiv:2608.23974v12026
  54. 6-DOF GraspNet: Variational Grasp Generation for Object Manipulation

    Arsalan Mousavian, Clemens Eppner, Dieter Fox

    cs.CVcs.ROarXiv:1905.10520v22019
  55. InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis

    Prateek Mittal, Ayush Srivastava, Joohi Chauhan

    q-bio.QMcs.CVcs.ITarXiv:2608.23574v12026
  56. Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

    Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le +10

    cs.CVcs.AIarXiv:2608.24130v12026
  57. Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

    Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen +10

    cs.CVcs.AIarXiv:2608.24154v12026
  58. Infant Care Video Dataset for Classification of Interventions Using Transformers

    Igor Bogdanov, James Green

    cs.CVcs.AIcs.LGarXiv:2608.23838v12026
  59. Communicating about Space: Language-Mediated Spatial Integration Across Partial Views

    Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil +2

    cs.CVarXiv:2603.27183v22026
  60. A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

    Kirill Skobelev, Eric Fithian, Yegor Baranovski +9

    cs.AIcs.CVcs.LGarXiv:2603.27341v42026