Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,881 to 14,940 of 18,868

  1. EfficientFormer: Vision Transformers at MobileNet Speed

    Yanyu Li, Geng Yuan, Yang Wen +5

    cs.CVarXiv:2206.01191v52022
  2. Deep learning with noisy labels: exploring techniques and remedies in medical image analysis

    Davood Karimi, Haoran Dou, Simon K. Warfield +1

    cs.CVcs.LGeess.IVarXiv:1912.02911v42019
  3. MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models

    Bocheng Zou, Mu Cai, Mark Stanley +2

    cs.CVarXiv:2603.25744v22026
  4. DeepStereo: Learning to Predict New Views from the World's Imagery

    John Flynn, Ivan Neulander, James Philbin +1

    cs.CVarXiv:1506.06825v12015
  5. Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification

    Dangwei Li, Xiaotang Chen, Zhang Zhang +1

    cs.CVarXiv:1710.06555v12017
  6. Unified Number-Free Text-to-Motion Generation Via Flow Matching

    Guanhe Huang, Oya Celiktutan

    cs.CVarXiv:2603.27040v12026
  7. CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets

    Robert N. Spaans, Catherine Chia, Tongjie Wang +7

    eess.IVcs.CVarXiv:2608.21497v12026
  8. Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

    Peixi Wu, Ke Mei, Feipeng Ma +15

    cs.CVarXiv:2604.22280v32026
    Summaries:한국어
  9. Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks

    Jie Hu, Li Shen, Samuel Albanie +2

    cs.CVarXiv:1810.12348v32018
  10. VirtualHome: Simulating Household Activities via Programs

    Xavier Puig, Kevin Ra, Marko Boben +4

    cs.CVcs.AIcs.LGarXiv:1806.07011v12018
  11. MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

    Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy +2

    cs.CVcs.AIarXiv:2608.24107v12026
  12. Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis

    Ming Cheng, Hongyu Sun, Zhaolin Chen +3

    cs.CVcs.MMarXiv:2608.23974v12026
  13. 6-DOF GraspNet: Variational Grasp Generation for Object Manipulation

    Arsalan Mousavian, Clemens Eppner, Dieter Fox

    cs.CVcs.ROarXiv:1905.10520v22019
  14. InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis

    Prateek Mittal, Ayush Srivastava, Joohi Chauhan

    q-bio.QMcs.CVcs.ITarXiv:2608.23574v12026
  15. Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking

    Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le +10

    cs.CVcs.AIarXiv:2608.24130v12026
  16. Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection

    Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen +10

    cs.CVcs.AIarXiv:2608.24154v12026
  17. Infant Care Video Dataset for Classification of Interventions Using Transformers

    Igor Bogdanov, James Green

    cs.CVcs.AIcs.LGarXiv:2608.23838v12026
  18. Communicating about Space: Language-Mediated Spatial Integration Across Partial Views

    Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil +2

    cs.CVarXiv:2603.27183v22026
  19. A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling

    Kirill Skobelev, Eric Fithian, Yegor Baranovski +9

    cs.AIcs.CVcs.LGarXiv:2603.27341v42026
  20. Towards Universal Fake Image Detectors that Generalize Across Generative Models

    Utkarsh Ojha, Yuheng Li, Yong Jae Lee

    cs.CVcs.LGarXiv:2302.10174v22023
  21. When and why vision-language models behave like bags-of-words, and what to do about it?

    Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri +2

    cs.CVcs.AIcs.CLarXiv:2210.01936v32022
  22. Falcon Perception

    Aviraj Bevli, Sofian Chaybouti, Yasser Dahou +6

    cs.CVarXiv:2603.27365v12026
  23. Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers

    Yuhe Liu, Zhenxiong Tan, Yujia Hu +2

    cs.CVarXiv:2603.27666v12026
  24. RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching

    Lahav Lipson, Zachary Teed, Jia Deng

    cs.CVarXiv:2109.07547v12021
  25. Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

    Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang +10

    cs.CVarXiv:2608.24138v12026
  26. CNN: Single-label to Multi-label

    Yunchao Wei, Wei Xia, Junshi Huang +4

    cs.CVarXiv:1406.5726v32014
  27. An Unsupervised Learning Model for Deformable Medical Image Registration

    Guha Balakrishnan, Amy Zhao, Mert R. Sabuncu +2

    cs.CVarXiv:1802.02604v32018
  28. A Survey of Deep Learning Applications to Autonomous Vehicle Control

    Sampo Kuutti, Richard Bowden, Yaochu Jin +2

    cs.LGcs.CVeess.SYarXiv:1912.10773v12019
  29. Deep Sliding Shapes for Amodal 3D Object Detection in RGB-D Images

    Shuran Song, Jianxiong Xiao

    cs.CVarXiv:1511.02300v22015
  30. A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision

    Vasiliki Kondyli, Jakob Suchan, Mehul Bhatt

    q-bio.NCcs.AIcs.CVarXiv:2608.23572v12026
  31. Fast-SCNN: Fast Semantic Segmentation Network

    Rudra P K Poudel, Stephan Liwicki, Roberto Cipolla

    cs.CVarXiv:1902.04502v12019
  32. LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

    Haoning Wu, Dongxu Li, Bei Chen +1

    cs.CVcs.CLcs.LGarXiv:2407.15754v12024
  33. PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models

    Lorenza Prospero, Orest Kupyn, Ostap Viniavskyi +2

    cs.CVarXiv:2603.28763v12026
  34. Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning

    Tommaso Galliena, Stefano Rosa, Tommaso Apicella +3

    cs.CVarXiv:2603.24257v22026
  35. SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering

    Wenli Li, Kai Zhao, Haoran Jiang +3

    cs.CVarXiv:2603.29437v12026
  36. Asymmetric Non-local Neural Networks for Semantic Segmentation

    Zhen Zhu, Mengde Xu, Song Bai +2

    cs.CVcs.LGarXiv:1908.07678v52019
  37. ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

    Yuanhao Sun, Huawei Ji, Jiaxin Ding +2

    cs.CVarXiv:2608.22996v12026
  38. Robust and Efficient Subspace Segmentation via Least Squares Regression

    Can-Yi Lu, Hai Min, Zhong-Qiu Zhao +3

    cs.CVarXiv:1404.6736v12014
  39. CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

    Oier Mees, Lukas Hermann, Erick Rosete-Beas +1

    cs.ROcs.AIcs.CLarXiv:2112.03227v42021
  40. LERF: Language Embedded Radiance Fields

    Justin Kerr, Chung Min Kim, Ken Goldberg +2

    cs.CVcs.GRarXiv:2303.09553v12023
  41. TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering

    Yunseok Jang, Yale Song, Youngjae Yu +2

    cs.CVarXiv:1704.04497v32017
  42. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation

    Bharath Krishnamurthy, Ajita Rattani

    cs.CVcs.AIarXiv:2603.29029v12026
  43. NeRV: Neural Reflectance and Visibility Fields for Relighting and View Synthesis

    Pratul P. Srinivasan, Boyang Deng, Xiuming Zhang +3

    cs.CVcs.GRarXiv:2012.03927v12020
  44. StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets

    Axel Sauer, Katja Schwarz, Andreas Geiger

    cs.LGcs.CVarXiv:2202.00273v22022
  45. Contrastive learning of global and local features for medical image segmentation with limited annotations

    Krishna Chaitanya, Ertunc Erdil, Neerav Karani +1

    cs.CVcs.LGeess.IVarXiv:2006.10511v22020
  46. RawGen: Learning Camera Raw Image Generation

    Dongyoung Kim, Junyong Lee, Abhijith Punnappurath +4

    cs.CVarXiv:2604.00093v12026
  47. Improving Diffusion Models for Inverse Problems using Manifold Constraints

    Hyungjin Chung, Byeongsu Sim, Dohoon Ryu +1

    cs.LGcs.AIcs.CVarXiv:2206.00941v32022
  48. HP-UniIF: Hierarchical Prompt Learning for Unified Image Fusion

    Xingxin Xu, Siqi Zhao, Xin Li +3

    cs.CVarXiv:2608.21786v12026
  49. Which Tasks Should Be Learned Together in Multi-task Learning?

    Trevor Standley, Amir R. Zamir, Dawn Chen +3

    cs.CVarXiv:1905.07553v42019
  50. PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search

    Yuhui Xu, Lingxi Xie, Xiaopeng Zhang +4

    cs.CVcs.LGarXiv:1907.05737v42019
  51. Learning by Cheating

    Dian Chen, Brady Zhou, Vladlen Koltun +1

    cs.ROcs.AIcs.CVarXiv:1912.12294v12019
  52. Measuring Robustness to Natural Distribution Shifts in Image Classification

    Rohan Taori, Achal Dave, Vaishaal Shankar +3

    cs.LGcs.CVstat.MLarXiv:2007.00644v22020
  53. SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers

    Nanye Ma, Mark Goldstein, Michael S. Albergo +3

    cs.CVcs.LGarXiv:2401.08740v22024
  54. Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion

    Hongxu Yin, Pavlo Molchanov, Zhizhong Li +5

    cs.LGcs.CVstat.MLarXiv:1912.08795v22019
  55. InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs

    Yujun Shen, Ceyuan Yang, Xiaoou Tang +1

    cs.CVcs.LGeess.IVarXiv:2005.09635v22020
  56. ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology

    Duncan Stothers, Ren-Chin Wu, William Lotter

    cs.CVcs.AIcs.LGarXiv:2608.22066v12026
  57. Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection

    Debidatta Dwibedi, Ishan Misra, Martial Hebert

    cs.CVarXiv:1708.01642v12017
  58. Gate Voltage Effect on Pulse Detection Efficiency of Perimeter-Gated SPADs

    Hunter Guthrie, Md Sakibur Sajal, Zexi Liu +1

    physics.ins-detcs.CVarXiv:2608.21371v12026
  59. VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

    Wenhai Wang, Zhe Chen, Xiaokang Chen +8

    cs.CVarXiv:2305.11175v22023
  60. Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label

    Cheng Chen, Yifan Zhao, Jia Li

    cs.CVarXiv:2608.22313v12026