Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

241 to 300 of 18,815

  1. Subject-driven Text-to-Image Generation via Apprenticeship Learning

    Wenhu Chen, Hexiang Hu, Yandong Li +4

    cs.CVcs.AIarXiv:2304.00186v52023
  2. Sparse MoEs meet Efficient Ensembles

    James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10

    cs.LGcs.CVstat.MLarXiv:2110.03360v22021
  3. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

    Zhengyi Wang, Cheng Lu, Yikai Wang +4

    cs.LGcs.CVarXiv:2305.16213v22023
  4. Cones: Concept Neurons in Diffusion Models for Customized Generation

    Zhiheng Liu, Ruili Feng, Kai Zhu +6

    cs.CVarXiv:2303.05125v12023
  5. Beyond neural scaling laws: beating power law scaling via data pruning

    Ben Sorscher, Robert Geirhos, Shashank Shekhar +2

    cs.LGcs.AIcs.CVarXiv:2206.14486v62022
  6. DaViT: Dual Attention Vision Transformers

    Mingyu Ding, Bin Xiao, Noel Codella +3

    cs.CVarXiv:2204.03645v12022
  7. Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

    A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1

    cs.CVcs.AIcs.LGarXiv:2604.18572v22026
  8. Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

    Lei Zhang, Junjiao Tian, Zhipeng Fan +9

    cs.CVarXiv:2604.04746v32026
  9. A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects

    Zewen Li, Wenjie Yang, Shouheng Peng +1

    cs.CVcs.LGeess.IVarXiv:2004.02806v12020
  10. SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

    Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman +12

    cs.CVcs.CEcs.CLarXiv:2602.03916v32026
  11. Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

    Anh Nguyen, Ngan Nguyen, Duc Vu +11

    cs.CVarXiv:2606.32020v12026
  12. Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

    Hohin Kwan, Hongyu Li, Ray Zhang +5

    cs.CVarXiv:2606.27828v12026
  13. SANet: Structure-Aware Network for Visual Tracking

    Heng Fan, Haibin Ling

    cs.CVarXiv:1611.06878v32016
  14. Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data

    Spencer Whitehead, Jacob Phillips, Sean Hendryx

    cs.CLcs.CVarXiv:2409.00238v12024
  15. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    Mantas Mazeika, Xuwang Yin, Rishub Tamirisa +8

    cs.LGcs.AIcs.CLarXiv:2502.08640v22025
  16. Efficient-CapsNet: Capsule Network with Self-Attention Routing

    Vittorio Mazzia, Francesco Salvetti, Marcello Chiaberge

    cs.CVcs.AIarXiv:2101.12491v22021
  17. Nerfies: Deformable Neural Radiance Fields

    Keunhong Park, Utkarsh Sinha, Jonathan T. Barron +4

    cs.CVcs.GRarXiv:2011.12948v52020
  18. AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection

    Yunkang Cao, Jiangning Zhang, Luca Frittoli +3

    cs.CVarXiv:2407.15795v12024
  19. FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image Classification

    Zhuo Zheng, Yanfei Zhong, Ailong Ma +1

    cs.CVeess.IVarXiv:2011.05670v12020
  20. Propagating Confidences through CNNs for Sparse Data Regression

    Abdelrahman Eldesokey, Michael Felsberg, Fahad Shahbaz Khan

    cs.CVcs.LGarXiv:1805.11913v32018
  21. Action Recognition with Image Based CNN Features

    Mahdyar Ravanbakhsh, Hossein Mousavi, Mohammad Rastegari +2

    cs.CVarXiv:1512.03980v12015
  22. DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model

    Siwei Xia, Li Sun, Tiantian Sun +1

    cs.CVarXiv:2505.12427v22025
  23. Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos

    Zhengze Xu, Mengting Chen, Zhao Wang +6

    cs.CVarXiv:2404.17571v12024
  24. 2023 Low-Power Computer Vision Challenge (LPCVC) Summary

    Leo Chen, Benjamin Boardley, Ping Hu +27

    cs.CVarXiv:2403.07153v12024
  25. Orca: The World is in Your Mind

    Yihao Wang, Yuheng Ji, Mingyu Cao +54

    cs.CVarXiv:2606.30534v32026
  26. Is the deconvolution layer the same as a convolutional layer?

    Wenzhe Shi, Jose Caballero, Lucas Theis +4

    cs.CVarXiv:1609.07009v12016
  27. End-to-end Lane Detection through Differentiable Least-Squares Fitting

    Wouter Van Gansbeke, Bert De Brabandere, Davy Neven +2

    cs.CVarXiv:1902.00293v32019
  28. Fruit Quality and Defect Image Classification with Conditional GAN Data Augmentation

    Jordan J. Bird, Chloe M. Barnes, Luis J. Manso +2

    cs.CVcs.LGeess.IVarXiv:2104.05647v12021
  29. Learning from Multimodal and Multitemporal Earth Observation Data for Building Damage Mapping

    Bruno Adriano, Naoto Yokoya, Junshi Xia +4

    cs.CVcs.LGarXiv:2009.06200v12020
  30. From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

    Zahra Anvari, Vassilis Athitsos

    cs.CLcs.CVarXiv:2609.17538v12026
  31. PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation

    Bo-Han Chen, Hiromu Taketsugu, Norimichi Ukita

    cs.CVcs.ROarXiv:2609.18125v12026
  32. A Multi-Stage model based on YOLOv3 for defect detection in PV panels based on IR and Visible Imaging by Unmanned Aerial Vehicle

    Antonio Di Tommaso, Alessandro Betti, Giacomo Fontanelli +1

    cs.CVcs.LGarXiv:2111.11709v22021
  33. TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning

    Yang Liu, Keze Wang, Lingbo Liu +2

    cs.CVarXiv:2112.03587v32021
  34. In-Context Robot Learning with VLM Agents

    Dongzhou Cheng, Taoran Yi, Ye Fang +12

    cs.CVcs.ROarXiv:2609.19138v12026
  35. FreqMamba: Viewing Mamba from a Frequency Perspective for Image Deraining

    Zou Zhen, Yu Hu, Zhao Feng

    cs.CVarXiv:2404.09476v22024
  36. Energy-Regularized Imitation Learning for Force- and Work-Aware Robotic Manipulation

    Toshiki Otani, Hiromu Taketsugu, Norimichi Ukita

    cs.ROcs.CVarXiv:2609.18164v12026
  37. Learning to Rank for Blind Image Quality Assessment

    Fei Gao, Dacheng Tao, Xinbo Gao +1

    cs.CVarXiv:1309.0213v32013
  38. VOIDD: automatic vessel of intervention dynamic detection in PCI procedures

    Ketan Bacchuwar, Jean Cousty, Régis Vaillant +1

    cs.CVarXiv:1710.04476v12017
  39. MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction

    Jung Min Lee, Dohyeok Lee, Seokhun Ju +5

    cs.ROcs.CVarXiv:2602.03668v32026
  40. OminiControl2: Efficient Conditioning for Diffusion Transformers

    Zhenxiong Tan, Qiaochu Xue, Xingyi Yang +2

    cs.CVcs.AIarXiv:2503.08280v12025
  41. Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception

    Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns

    cs.ROcs.CVcs.LGarXiv:2609.18100v12026
  42. FASA: Feature Augmentation and Sampling Adaptation for Long-Tailed Instance Segmentation

    Yuhang Zang, Chen Huang, Chen Change Loy

    cs.CVarXiv:2102.12867v22021
  43. Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

    Vishnu Bindu Balachandran

    cs.LGcs.AIcs.CVarXiv:2609.17560v12026
  44. MCLC-NET: Multimodal Continual Learning for Leaf Counting

    Ruchi Bhatt, Pratibha Kumari, Shreya Bansal +3

    cs.CVcs.LGarXiv:2609.18129v12026
  45. Temperon: Full-Time SAM Quality at a Third Less Wall-Clock

    Stamatis Mastromichalakis

    cs.LGcs.CVarXiv:2609.17575v12026
  46. 3D AGSE-VNet: An Automatic Brain Tumor MRI Data Segmentation Framework

    Xi Guan, Guang Yang, Jianming Ye +4

    cs.AIcs.CVcs.LGarXiv:2107.12046v12021
  47. HairCS: Reconstructing Strand-Based Hair from Hair Cards

    Zixuan Lu, Tongtong Wang, Yuefan Shen +4

    cs.GRcs.CVarXiv:2609.16465v12026
  48. Generative Video Motion Editing with 3D Point Tracks

    Yao-Chih Lee, Zhoutong Zhang, Jiahui Huang +5

    cs.CVarXiv:2512.02015v12025
  49. CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

    Yunxiang Fu, Meng Lou, Zicheng Liao +1

    cs.LGcs.CVarXiv:2609.17026v12026
  50. Different Approaches for Human Activity Recognition: A Survey

    Zawar Hussain, Michael Sheng, Wei Emma Zhang

    cs.CVarXiv:1906.05074v12019
  51. Less is More: Focus Attention for Efficient DETR

    Dehua Zheng, Wenhui Dong, Hailin Hu +2

    cs.CVcs.AIarXiv:2307.12612v12023
  52. A Comprehensive Review of Generative Physical Artificial Intelligence

    Satyam Gaba, Krutiksinh Rana, Siva Sai +2

    cs.ROcs.AIcs.CLarXiv:2609.18111v12026
  53. Data Augmentation for Skin Lesion using Self-Attention based Progressive Generative Adversarial Network

    Ibrahim Saad Ali, Mamdouh Farouk Mohamed, Yousef Bassyouni Mahdy

    eess.IVcs.CVarXiv:1910.11960v12019
  54. Rain rendering for evaluating and improving robustness to bad weather

    Maxime Tremblay, Shirsendu Sukanta Halder, Raoul de Charette +1

    cs.CVarXiv:2009.03683v12020
  55. Multimodal Engagement Analysis from Facial Videos in the Classroom

    Ömer Sümer, Patricia Goldberg, Sidney D'Mello +3

    cs.CVcs.MMarXiv:2101.04215v22021
  56. Bi-Level Routing and Sparse Spatial Attention based Multi-View BEV 3D Object Detection for Autonomous Driving

    Jing Zhang, Jiaqi Liu, Zibo Wang

    cs.CVcs.AIarXiv:2609.14185v12026
  57. Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

    Manglesh Kumar Pandey, Sumit Kumar Banshal

    cs.CVcs.LGarXiv:2609.16859v12026
  58. FRPSS: Feature Rearrangement in Pre-Shape Space for Single-Image Generation

    Yuexing Han, Haoxuan Zhang, Bing Wang

    cs.CVarXiv:2609.16594v12026
  59. HS-FPN: High Frequency and Spatial Perception FPN for Tiny Object Detection

    Zican Shi, Jing Hu, Jie Ren +6

    cs.CVarXiv:2412.10116v32024
  60. Adaptive AI: Energy Efficient Multi-exit TinyML on Intelligent Vision Systems at the Edge

    Luca Crupi, Lorenzo Lamberti, Alessandro Giusti +1

    cs.ARcs.CVcs.DCarXiv:2609.11939v12026