Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,301 to 3,360 of 18,822

  1. Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models

    Tolgay Atinc Uzun, Dmitry Ignatov, Radu Timofte

    cs.CVarXiv:2601.08517v22026
  2. Towards Privacy-Preserving Visual Recognition via Adversarial Training: A Pilot Study

    Zhenyu Wu, Zhangyang Wang, Zhaowen Wang +1

    cs.CVarXiv:1807.08379v22018
  3. SpinalNet: Deep Neural Network with Gradual Input

    H M Dipu Kabir, Moloud Abdar, Seyed Mohammad Jafar Jalali +4

    cs.CVcs.LGcs.NEarXiv:2007.03347v32020
  4. Training-Free Logical and Structural Anomaly Detection via Calibrated Fusion

    Changyi Li, Miao Yu, Kai Dong +1

    cs.CVarXiv:2609.05091v12026
  5. Cut to the Mix: Simple Data Augmentation Outperforms Elaborate Ones in Limited Organ Segmentation Datasets

    Chang Liu, Fuxin Fan, Annette Schwarz +1

    cs.CVarXiv:2602.03555v12026
  6. ReCo-KD: Region- and Context-Aware Knowledge Distillation for Efficient 3D Medical Image Segmentation

    Qizhen Lan, Yu-Chun Hsu, Nida Saddaf Khan +1

    cs.CVarXiv:2601.08301v12026
  7. MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

    Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah +2

    cs.CVarXiv:2609.05070v12026
  8. PuTR-CouT: Counting-by-Tracking in Camera-Trap Image Sequences

    Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos

    cs.CVarXiv:2609.05038v12026
  9. PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

    Xudong Lu, Huankang Guan, Yang Bo +10

    cs.CVcs.CLarXiv:2601.22575v12026
  10. Rethinking CNN Models for Audio Classification

    Kamalesh Palanisamy, Dipika Singhania, Angela Yao

    cs.CVcs.SDeess.ASarXiv:2007.11154v22020
  11. Measured Sliders: Learning Continuous Controls from Differentiable Image Measurements

    Yijia Chen, Boyu Wei, Xuanhua Yin

    cs.CVarXiv:2609.05234v12026
  12. First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

    Tianjie Ju, Xinyue Xu, Wanxuan Sun +4

    cs.CVarXiv:2609.05224v12026
  13. BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors

    Vincent Leroy, Philippe Weinzaepfel, Lojze Zust +2

    cs.CVarXiv:2609.05210v12026
  14. Densely Guided Knowledge Distillation using Multiple Teacher Assistants

    Wonchul Son, Jaemin Na, Junyong Choi +1

    cs.CVarXiv:2009.08825v32020
  15. WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

    Hui Zhang, Zongkai Liu, Liqiang Niu +6

    cs.CVarXiv:2609.05171v12026
  16. CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofing

    Ajian Li, Zichang Tan, Xuan Li +4

    cs.CVarXiv:2003.05136v12020
  17. VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

    Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer +5

    cs.CVarXiv:2609.05114v12026
  18. NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild

    Jason Y. Zhang, Gengshan Yang, Shubham Tulsiani +1

    cs.CVcs.LGarXiv:2110.07604v32021
  19. Trajectory Forecasts in Unknown Environments Conditioned on Grid-Based Plans

    Nachiket Deo, Mohan M. Trivedi

    cs.CVcs.ROarXiv:2001.00735v22020
  20. Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection

    Fredrik Lundell, Per-Erik Forssen, Mårten Wadenbäck +1

    cs.CVarXiv:2609.05049v12026
  21. OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models

    Xingyi He, Jiaming Sun, Yuang Wang +3

    cs.CVarXiv:2301.07673v12023
  22. Learning 3D Editing without Paired Supervision via Generative Prior Distillation

    Hao Wen, Weibin Yun, Hongxing Fan +4

    cs.CVarXiv:2609.04942v12026
  23. ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization

    Hyeongsik Kim, Mincheol Kim, Heejoon Moon +1

    cs.CVarXiv:2609.04965v12026
  24. MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

    Zijie Zhu, Weiren Cai, Yizhou Wang +4

    cs.CVcs.ROarXiv:2609.04958v12026
  25. TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image

    Xin Zhang, Yabo Chen, Zixuan Duan +4

    cs.CVarXiv:2609.04911v12026
  26. LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering

    Yachuan Huang, Liwen Xiao, Liao Shen +4

    cs.CVarXiv:2609.04939v12026
  27. Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation

    Jia Li, Xiaomeng Fu, Xurui Peng +7

    cs.CVarXiv:2602.14027v32026
  28. SASA: Semantics-Augmented Set Abstraction for Point-based 3D Object Detection

    Chen Chen, Zhe Chen, Jing Zhang +1

    cs.CVarXiv:2201.01976v12022
  29. InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

    Yihan Zhou, Zikai Huang, Yuyang Yu +3

    cs.CVarXiv:2609.04903v12026
  30. Domain Generalization via Shuffled Style Assembly for Face Anti-Spoofing

    Zhuo Wang, Zezheng Wang, Zitong Yu +4

    cs.CVarXiv:2203.05340v42022
  31. Compositional Reward Models for Conditional Medical Image Generation

    Aayush Kumar Tyagi, Prathosh A. P., Mausam

    cs.CVarXiv:2609.05028v12026
  32. Temporal Multimodal Fusion for Video Emotion Classification in the Wild

    Valentin Vielzeuf, Stéphane Pateux, Frédéric Jurie

    cs.CVcs.LGcs.MMarXiv:1709.07200v12017
  33. A Signal Propagation Perspective for Pruning Neural Networks at Initialization

    Namhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould +1

    cs.LGcs.CVstat.MLarXiv:1906.06307v22019
  34. RefDiT: Local Attribute Guidance in Reference-Based Image Generation

    Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam

    cs.CVarXiv:2609.04976v12026
  35. Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes

    Jiashun Wang, Huazhe Xu, Jingwei Xu +2

    cs.CVarXiv:2012.05522v22020
  36. Task Residual for Tuning Vision-Language Models

    Tao Yu, Zhihe Lu, Xin Jin +2

    cs.CVarXiv:2211.10277v22022
  37. Learning Fully Convolutional Networks for Iterative Non-blind Deconvolution

    Jiawei Zhang, Jinshan Pan, Wei-Sheng Lai +2

    cs.CVarXiv:1611.06495v12016
  38. Image Generators with Conditionally-Independent Pixel Synthesis

    Ivan Anokhin, Kirill Demochkin, Taras Khakhulin +3

    cs.CVcs.AIcs.LGarXiv:2011.13775v12020
  39. Only a Matter of Style: Age Transformation Using a Style-Based Regression Model

    Yuval Alaluf, Or Patashnik, Daniel Cohen-Or

    cs.CVarXiv:2102.02754v22021
  40. Compact Generalized Non-local Network

    Kaiyu Yue, Ming Sun, Yuchen Yuan +3

    cs.CVarXiv:1810.13125v22018
  41. LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus

    Chi Zhang, Qi Song, Feifei Li +2

    cs.CVarXiv:2609.04846v12026
  42. Periodic Vibration Gaussian: Dynamic Urban Scene Reconstruction and Real-time Rendering

    Yurui Chen, Chun Gu, Junzhe Jiang +2

    cs.CVarXiv:2311.18561v32023
  43. Semi-Supervised Learning of Visual Features by Non-Parametrically Predicting View Assignments with Support Samples

    Mahmoud Assran, Mathilde Caron, Ishan Misra +4

    cs.CVcs.AIcs.LGarXiv:2104.13963v32021
  44. Weather-Conditioned Depth Anything

    Zhaoming Xu, Chan-Wei Hu, Kuan-Ru Huang +4

    cs.CVarXiv:2609.04827v12026
  45. LUMIN: Lightweight Universal Manufacturing Inspection Network for Anomaly Detection

    Pengfei Yang

    cs.CVarXiv:2609.04775v12026
  46. Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency Coherence

    Siyue Yu, Bingfeng Zhang, Jimin Xiao +1

    cs.CVarXiv:2012.04404v22020
  47. CoMLP: Cooperatively-Gated MLPs for Fine-Grained Cross-Modal Information Fusion in Medical Image Segmentation

    Mingyuan Meng, Shuchang Ye, Mingjian Li +3

    cs.CVarXiv:2609.04781v12026
  48. SeamFlow: Structure-Aware Flow Matching on Edge Probabilities for Artist-Like UV Unwrapping

    Yuming Zhao, Zangyueyang Xian, Qijian Zhang +4

    cs.CVarXiv:2609.04751v12026
  49. Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework

    Yibo Yan, Mingdong Ou, Yi Cao +5

    cs.CLcs.CVcs.IRarXiv:2602.19549v22026
  50. Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

    Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu +4

    cs.CVarXiv:2609.04741v12026
  51. HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction

    Ziding Zhao, Hainan Cui, Peilin Tao +1

    cs.CVarXiv:2609.04718v12026
  52. The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

    Yizhak Ben-Shabat, Xin Yu, Fatemeh Sadat Saleh +4

    cs.CVarXiv:2007.00394v22020
  53. FCAF3D: Fully Convolutional Anchor-Free 3D Object Detection

    Danila Rukhovich, Anna Vorontsova, Anton Konushin

    cs.CVarXiv:2112.00322v22021
  54. Weakly-Supervised Video Moment Retrieval via Semantic Completion Network

    Zhijie Lin, Zhou Zhao, Zhu Zhang +2

    cs.CVcs.LGcs.MMarXiv:1911.08199v32019
  55. Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

    Hyun Seok Seong, Woojin Jun, SuBeen Lee +1

    cs.CVarXiv:2609.04800v12026
  56. From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

    Usha Shrestha, Dmitry Ignatov, Radu Timofte

    cs.CVcs.LGarXiv:2601.03808v22026
  57. Large-Scale Visual Relationship Understanding

    Ji Zhang, Yannis Kalantidis, Marcus Rohrbach +3

    cs.CVarXiv:1804.10660v42018
  58. An Attention-Guided Global and Local Fusion Framework for Lesion-Focused Image Classification

    Mst Shafia Tasnima, Md Samaun Elaheea, Tanjim Taharat Aurpab +1

    cs.CVarXiv:2609.04791v12026
  59. AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing

    Jiaxi Jiang, Paul Streli, Huajian Qiu +4

    cs.CVcs.AIcs.GRarXiv:2207.13784v12022
  60. AdaPoinTr: Diverse Point Cloud Completion with Adaptive Geometry-Aware Transformers

    Xumin Yu, Yongming Rao, Ziyi Wang +2

    cs.CVcs.AIarXiv:2301.04545v12023