Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,821 to 14,880 of 18,849
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Ke Wang, Junting Pan, Weikang Shi +3
cs.CVcs.AIcs.CLarXiv:2402.14804v12024Stereo World Model: Camera-Guided Stereo Video Generation
Yang-Tian Sun, Zehuan Huang, Yifan Niu +4
cs.CVarXiv:2603.17375v12026VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi +3
cs.CVcs.AIarXiv:2603.17948v120263DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
Hyun-kyu Ko, Jihyeon Park, Younghyun Kim +2
cs.CVarXiv:2603.18524v12026In-the-Wild Camouflage Attack on Vehicle Detectors through Controllable Image Editing
Xiao Fang, Yiming Gong, Stanislav Panev +4
cs.CVarXiv:2603.19456v12026LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation
Junjue Wang, Zhuo Zheng, Ailong Ma +2
cs.CVarXiv:2110.08733v62021TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos
Yan Zeng, Haoran Jiang, Kaixin Yao +4
cs.CVarXiv:2603.17735v12026Benchmarking Denoising Algorithms with Real Photographs
Tobias Plötz, Stefan Roth
cs.CVarXiv:1707.01313v12017Context-Aware Crowd Counting
Weizhe Liu, Mathieu Salzmann, Pascal Fua
cs.CVarXiv:1811.10452v22018Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Yixin Liu, Kai Zhang, Yuan Li +9
cs.CVcs.AIcs.LGarXiv:2402.17177v32024Robust Global Structure-from-Motion via View Graph Pruning
Jiamin Xu, Lixing Yao, Weichen Dai +4
cs.CVarXiv:2608.22054v12026BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
Sheng Zhang, Yanbo Xu, Naoto Usuyama +21
cs.CVcs.CLarXiv:2303.00915v32023Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
Shang-Jui Ray Kuo, Paola Cascante-Bonilla
cs.CVcs.LGarXiv:2603.19209v12026Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
Youbin Kim, Jinho Park, Hogun Park +1
cs.CVarXiv:2603.21944v12026Generalized ODIN: Detecting Out-of-distribution Image without Learning from Out-of-distribution Data
Yen-Chang Hsu, Yilin Shen, Hongxia Jin +1
cs.CVcs.LGeess.IVarXiv:2002.11297v22020Event-Based Motion Estimation via Oriented Distance Fields
Lei Sun, Yuqin Ma, Weilun Li +5
cs.CVcs.ROarXiv:2608.24223v12026VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas +2
cs.CVcs.AIcs.LGarXiv:2603.23495v12026Scaling up GANs for Text-to-Image Synthesis
Minguk Kang, Jun-Yan Zhu, Richard Zhang +4
cs.CVcs.GRcs.LGarXiv:2303.05511v22023Unified Focal loss: Generalising Dice and cross entropy-based losses to handle class imbalanced medical image segmentation
Michael Yeung, Evis Sala, Carola-Bibiane Schönlieb +1
eess.IVcs.CVcs.LGarXiv:2102.04525v42021F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
Injae Kim, Chaehyeon Kim, Minseong Bae +2
cs.CVarXiv:2603.21304v22026Scaling Out-of-Distribution Detection for Real-World Settings
Dan Hendrycks, Steven Basart, Mantas Mazeika +5
cs.CVcs.LGarXiv:1911.11132v42019Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR
Sheethal Bhat, Bogdan Georgescu, Awais Mansoor +5
cs.CVarXiv:2608.24281v12026Generalized Zero- and Few-Shot Learning via Aligned Variational Autoencoders
Edgar Schönfeld, Sayna Ebrahimi, Samarth Sinha +2
cs.CVcs.AIcs.LGarXiv:1812.01784v420184DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
Jae Won Jang, Yeonjin Chang, Wonsik Shin +2
cs.CVarXiv:2603.21618v22026Revisiting Batch Normalization For Practical Domain Adaptation
Yanghao Li, Naiyan Wang, Jianping Shi +2
cs.CVcs.LGarXiv:1603.04779v42016Robotic Pick-and-Place of Novel Objects in Clutter with Multi-Affordance Grasping and Cross-Domain Image Matching
Andy Zeng, Shuran Song, Kuan-Ting Yu +18
cs.ROcs.CVarXiv:1710.01330v52017TrajLoom: Dense Future Trajectory Generation from Video
Zewei Zhang, Jia Jun Cheng Xian, Kaiwen Liu +4
cs.CVarXiv:2603.22606v12026BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
Risa Shinoda, Kaede Shiohara, Nakamasa Inoue +3
cs.CVarXiv:2603.23883v12026RealMaster: Lifting Rendered Scenes into Photorealistic Video
Dana Cohen-Bar, Ido Sobol, Raphael Bensadoun +5
cs.CVarXiv:2603.23462v12026DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models
Jaewon Min, Jaeeun Lee, Yeji Choi +7
cs.CVarXiv:2603.23499v12026Mobile-Former: Bridging MobileNet and Transformer
Yinpeng Chen, Xiyang Dai, Dongdong Chen +4
cs.CVcs.LGarXiv:2108.05895v32021Non-Local Recurrent Network for Image Restoration
Ding Liu, Bihan Wen, Yuchen Fan +2
cs.CVarXiv:1806.02919v22018PixelSmile: Toward Fine-Grained Facial Expression Editing
Jiabin Hua, Hengyuan Xu, Aojie Li +4
cs.CVcs.AIarXiv:2603.25728v12026TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans - A Case Study in Craniosynostosis 3D Photography
Sanjay Bhandari, Nawazish Khan, Alzbeta Novotna +7
cs.CVcs.AIarXiv:2608.22131v12026Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
Ruixing Zhang, Hanzhang Jiang, Leilei Sun +3
cs.CVcs.AIarXiv:2603.26610v12026One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
Adrien Ramanana Rahary, Nicolas Dufour, Patrick Perez +1
cs.CVarXiv:2603.23488v22026Structural Graph Probing of Vision-Language Models
Haoyu He, Yue Zhuo, Yu Zheng +1
cs.CVarXiv:2603.27070v12026SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision
Avigail Cohen Rimon, Amir Mann, Mirela Ben Chen +1
cs.CVarXiv:2603.24036v22026DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery
Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas +1
cs.CVcs.CLcs.IRarXiv:2608.21529v12026Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
Danil Tokhchukov, Aysel Mirzoeva, Andrey Kuznetsov +1
cs.CVarXiv:2603.24800v22026Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6
cs.AIcs.CLcs.CVarXiv:2603.24961v12026EfficientFormer: Vision Transformers at MobileNet Speed
Yanyu Li, Geng Yuan, Yang Wen +5
cs.CVarXiv:2206.01191v52022Deep learning with noisy labels: exploring techniques and remedies in medical image analysis
Davood Karimi, Haoran Dou, Simon K. Warfield +1
cs.CVcs.LGeess.IVarXiv:1912.02911v42019MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Bocheng Zou, Mu Cai, Mark Stanley +2
cs.CVarXiv:2603.25744v22026DeepStereo: Learning to Predict New Views from the World's Imagery
John Flynn, Ivan Neulander, James Philbin +1
cs.CVarXiv:1506.06825v12015Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
Dangwei Li, Xiaotang Chen, Zhang Zhang +1
cs.CVarXiv:1710.06555v12017Unified Number-Free Text-to-Motion Generation Via Flow Matching
Guanhe Huang, Oya Celiktutan
cs.CVarXiv:2603.27040v12026CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets
Robert N. Spaans, Catherine Chia, Tongjie Wang +7
eess.IVcs.CVarXiv:2608.21497v12026Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
cs.CVarXiv:2604.22280v32026Summaries:한국어Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
Jie Hu, Li Shen, Samuel Albanie +2
cs.CVarXiv:1810.12348v32018VirtualHome: Simulating Household Activities via Programs
Xavier Puig, Kevin Ra, Marko Boben +4
cs.CVcs.AIcs.LGarXiv:1806.07011v12018MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes
Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy +2
cs.CVcs.AIarXiv:2608.24107v12026Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis
Ming Cheng, Hongyu Sun, Zhaolin Chen +3
cs.CVcs.MMarXiv:2608.23974v120266-DOF GraspNet: Variational Grasp Generation for Object Manipulation
Arsalan Mousavian, Clemens Eppner, Dieter Fox
cs.CVcs.ROarXiv:1905.10520v22019InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis
Prateek Mittal, Ayush Srivastava, Joohi Chauhan
q-bio.QMcs.CVcs.ITarXiv:2608.23574v12026Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking
Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le +10
cs.CVcs.AIarXiv:2608.24130v12026Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection
Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen +10
cs.CVcs.AIarXiv:2608.24154v12026Infant Care Video Dataset for Classification of Interventions Using Transformers
Igor Bogdanov, James Green
cs.CVcs.AIcs.LGarXiv:2608.23838v12026Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil +2
cs.CVarXiv:2603.27183v22026A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
Kirill Skobelev, Eric Fithian, Yegor Baranovski +9
cs.AIcs.CVcs.LGarXiv:2603.27341v42026