Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,941 to 15,000 of 18,867

  1. RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

    Oguzhan Baser, Mirac Sozen, Kaan Kale +2

    cs.ROcs.AIcs.CVarXiv:2608.21380v12026
  2. Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

    Rui Xia, Ayan Das, Artem Artemev +3

    cs.CVstat.MLarXiv:2608.21972v12026
  3. Recurrent MVSNet for High-resolution Multi-view Stereo Depth Inference

    Yao Yao, Zixin Luo, Shiwei Li +3

    cs.CVarXiv:1902.10556v12019
  4. Differentiable Voxelization of Surface Representations

    Tobias Djuren, Ugo Finnendahl, Markus Worchel +2

    cs.GRcs.CVarXiv:2608.15934v12026
  5. PointNetVLAD: Deep Point Cloud Based Retrieval for Large-Scale Place Recognition

    Mikaela Angelina Uy, Gim Hee Lee

    cs.CVarXiv:1804.03492v32018
  6. 3D Point Cloud from Close-Range Photogrammetry for Defect Characterisation of Rubberised Concrete

    Jiacheng Liu, Mohammed Alnahhal, Ailar Hajimohammadi +3

    cs.CVarXiv:2608.21468v12026
  7. Executing your Commands via Motion Diffusion in Latent Space

    Xin Chen, Biao Jiang, Wen Liu +5

    cs.CVcs.GRarXiv:2212.04048v32022
  8. Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging

    Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto +3

    cs.CVarXiv:2608.21970v12026
  9. On the Choice of Tensor Estimation for Corner Detection, Optical Flow and Denoising

    Freddie Åström, Michael Felsberg

    cs.CVarXiv:2608.22314v12026
  10. Searching for Activation Functions

    Prajit Ramachandran, Barret Zoph, Quoc V. Le

    cs.NEcs.CVcs.LGarXiv:1710.05941v22017
  11. MedGAN: Medical Image Translation using GANs

    Karim Armanious, Chenming Jiang, Marc Fischer +4

    cs.CVarXiv:1806.06397v22018
  12. Dual Attention Networks for Multimodal Reasoning and Matching

    Hyeonseob Nam, Jung-Woo Ha, Jeonghee Kim

    cs.CVarXiv:1611.00471v22016
  13. V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer

    Runsheng Xu, Hao Xiang, Zhengzhong Tu +3

    cs.CVarXiv:2203.10638v32022
  14. Unsupervised Deep Feature Extraction for Remote Sensing Image Classification

    Adriana Romero, Carlo Gatta, Gustau Camps-Valls

    cs.CVarXiv:1511.08131v12015
  15. MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices

    Sheng Chen, Yang Liu, Xiang Gao +1

    cs.CVcs.LGarXiv:1804.07573v42018
  16. Targeted Iterative Filtering

    Freddie Åström, Michael Felsberg, George Baravdish +1

    cs.CVarXiv:2608.22299v12026
  17. Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

    Boseung Jeong, Taegyu Park, Donghyeon Kwon +2

    cs.CVarXiv:2608.22500v12026
  18. On Tensor-Based PDEs and their Corresponding Variational Formulations with Application to Color Image Denoising

    Freddie Åström, George Baravdish, Michael Felsberg

    cs.CVarXiv:2608.22302v12026
  19. LiST: Local-Simplex Test-Time LoRA Fusion

    Yihua Shao, Jia Li, Siyu Chen +12

    cs.CVarXiv:2608.22370v12026
  20. Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition

    Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh

    cs.CVarXiv:1708.07632v12017
  21. Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models

    Yuanyang Yin, Yufan Deng, Shenghai Yuan +3

    cs.CVarXiv:2601.07287v12026
  22. What makes ImageNet good for transfer learning?

    Minyoung Huh, Pulkit Agrawal, Alexei A. Efros

    cs.CVcs.AIcs.LGarXiv:1608.08614v22016
  23. No Fuss Distance Metric Learning using Proxies

    Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung +2

    cs.CVarXiv:1703.07464v32017
  24. SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL

    Lijun Liu, Linwei Chen, Zhishou Zhang +7

    cs.CVcs.AIarXiv:2601.09136v12026
  25. mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

    Qinghao Ye, Haiyang Xu, Jiabo Ye +7

    cs.CLcs.CVarXiv:2311.04257v22023
  26. CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers

    Jiaming Zhang, Huayao Liu, Kailun Yang +3

    cs.CVcs.ROeess.IVarXiv:2203.04838v52022
  27. Exploring the Limitations of Behavior Cloning for Autonomous Driving

    Felipe Codevilla, Eder Santana, Antonio M. López +1

    cs.CVcs.AIarXiv:1904.08980v12019
  28. Bee Detection and Tracking at Hive Entrance using YOLO11 and ByteTrack

    Thi Thu Thao Nguyen, Johannes Reschke

    cs.CVarXiv:2608.23213v12026
  29. Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

    Dohun Lee, Chun-Hao Paul Huang, Xuelin Chen +3

    cs.CVcs.AIcs.LGarXiv:2601.16296v22026
  30. DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

    Hengyu Shen, Tiancheng Gu, Bin Qin +10

    cs.CVcs.AIarXiv:2601.10305v32026
  31. 3D Shape Generation and Completion through Point-Voxel Diffusion

    Linqi Zhou, Yilun Du, Jiajun Wu

    cs.CVarXiv:2104.03670v32021
  32. FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects

    Bowen Wen, Wei Yang, Jan Kautz +1

    cs.CVcs.AIcs.ROarXiv:2312.08344v22023
  33. MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing

    Kai Zhang, Lingbo Mo, Wenhu Chen +2

    cs.CVcs.AIcs.CLarXiv:2306.10012v32023
  34. Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing

    Tingyu Song, Yanzhao Zhang, Mingxin Li +6

    cs.CVcs.CLcs.IRarXiv:2601.16125v12026
  35. O-HAZE: a dehazing benchmark with real hazy and haze-free outdoor images

    Codruta O. Ancuti, Cosmin Ancuti, Radu Timofte +1

    cs.CVarXiv:1804.05101v12018
  36. Structured Knowledge Distillation for Dense Prediction

    Yifan Liu, Changyong Shun, Jingdong Wang +1

    cs.CVarXiv:1903.04197v72019
  37. Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

    Xiwei Liu, Yulong Li, Xinlin Zhuang +5

    cs.CVarXiv:2608.23074v12026
  38. Domain Generalization for Object Recognition with Multi-task Autoencoders

    Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang +1

    cs.CVcs.AIcs.LGarXiv:1508.07680v12015
  39. Following Motion for Sequential Modeling in Video Frame Interpolation

    Jaehyun Park, Nam Ik Cho

    cs.CVarXiv:2608.22861v12026
  40. Artificial Intelligence in the Creative Industries: A Review

    Nantheera Anantrasirichai, David Bull

    cs.CVcs.AIcs.LGarXiv:2007.12391v62020
  41. TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation

    Ariel Shaulov, Eitan Shaar, Amit Edenzon +1

    cs.CVcs.AIarXiv:2602.00268v12026
  42. Ten Years of Pedestrian Detection, What Have We Learned?

    Rodrigo Benenson, Mohamed Omran, Jan Hosang +1

    cs.CVarXiv:1411.4304v12014
  43. Rotation-invariant convolutional neural networks for galaxy morphology prediction

    Sander Dieleman, Kyle W. Willett, Joni Dambre

    astro-ph.IMastro-ph.GAcs.CVarXiv:1503.07077v12015
  44. Seeing the Unseen: Semantic-in-Gaussian for Sparse-View 3D Generalization

    Zeyang Bai, Yunpeng Wang, Yunbiao Wang +1

    cs.CVarXiv:2608.22740v12026
  45. Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks

    Meng-Hao Guo, Zheng-Ning Liu, Tai-Jiang Mu +1

    cs.CVarXiv:2105.02358v22021
  46. Deep Learning in Multimodal Remote Sensing Data Fusion: A Comprehensive Review

    Jiaxin Li, Danfeng Hong, Lianru Gao +4

    cs.CVcs.LGeess.SParXiv:2205.01380v12022
  47. Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data

    Grzegorz Stefanski, Alberto Presta, Michal Byra

    cs.AIcs.CVcs.LGarXiv:2601.22141v12026
  48. AI-Generated Image Detectors Overrely on Global Artifacts: Evidence from Inpainting Exchange

    Elif Nebioglu, Emirhan Bilgiç, Adrian Popescu

    cs.CVcs.AIarXiv:2602.00192v12026
  49. DSAC - Differentiable RANSAC for Camera Localization

    Eric Brachmann, Alexander Krull, Sebastian Nowozin +4

    cs.CVarXiv:1611.05705v42016
  50. Searching for A Robust Neural Architecture in Four GPU Hours

    Xuanyi Dong, Yi Yang

    cs.CVarXiv:1910.04465v22019
  51. Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding

    Kexin Yi, Jiajun Wu, Chuang Gan +3

    cs.AIcs.CLcs.CVarXiv:1810.02338v22018
  52. Reliable and Responsible Foundation Models: A Comprehensive Survey

    Xinyu Yang, Junlin Han, Rishi Bommasani +49

    cs.LGcs.AIcs.CLarXiv:2602.08145v12026
  53. PlantDoc: A Dataset for Visual Plant Disease Detection

    Davinder Singh, Naman Jain, Pranjali Jain +3

    cs.CVeess.IVarXiv:1911.10317v12019
  54. Contextualized Visual Personalization in Vision-Language Models

    Yeongtak Oh, Sangwon Yu, Junsung Park +3

    cs.CVarXiv:2602.03454v32026
  55. Concealed Object Detection

    Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng +1

    cs.CVarXiv:2102.10274v22021
  56. POP: Prefill-Only Pruning for Efficient Large Model Inference

    Junhui He, Zhihui Fu, Jun Wang +1

    cs.CLcs.AIcs.CVarXiv:2602.03295v22026
  57. Aggregating Deep Convolutional Features for Image Retrieval

    Artem Babenko, Victor Lempitsky

    cs.CVarXiv:1510.07493v12015
  58. FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation

    Zimu Lu, Houxing Ren, Yunqiao Yang +4

    cs.SEcs.CLcs.CVarXiv:2602.03798v12026
  59. In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

    Behnam Neyshabur, Ryota Tomioka, Nathan Srebro

    cs.LGcs.AIcs.CVarXiv:1412.6614v42014
  60. High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks

    Haohan Wang, Xindi Wu, Zeyi Huang +1

    cs.CVcs.LGarXiv:1905.13545v32019