Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
241 to 300 of 18,815
Subject-driven Text-to-Image Generation via Apprenticeship Learning
Wenhu Chen, Hexiang Hu, Yandong Li +4
cs.CVcs.AIarXiv:2304.00186v52023Sparse MoEs meet Efficient Ensembles
James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10
cs.LGcs.CVstat.MLarXiv:2110.03360v22021ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
Zhengyi Wang, Cheng Lu, Yikai Wang +4
cs.LGcs.CVarXiv:2305.16213v22023Cones: Concept Neurons in Diffusion Models for Customized Generation
Zhiheng Liu, Ruili Feng, Kai Zhu +6
cs.CVarXiv:2303.05125v12023Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar +2
cs.LGcs.AIcs.CVarXiv:2206.14486v62022DaViT: Dual Attention Vision Transformers
Mingyu Ding, Bin Xiao, Noel Codella +3
cs.CVarXiv:2204.03645v12022Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1
cs.CVcs.AIcs.LGarXiv:2604.18572v22026Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Lei Zhang, Junjiao Tian, Zhipeng Fan +9
cs.CVarXiv:2604.04746v32026A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
Zewen Li, Wenjie Yang, Shouheng Peng +1
cs.CVcs.LGeess.IVarXiv:2004.02806v12020SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman +12
cs.CVcs.CEcs.CLarXiv:2602.03916v32026Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Anh Nguyen, Ngan Nguyen, Duc Vu +11
cs.CVarXiv:2606.32020v12026Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Hohin Kwan, Hongyu Li, Ray Zhang +5
cs.CVarXiv:2606.27828v12026SANet: Structure-Aware Network for Visual Tracking
Heng Fan, Haibin Ling
cs.CVarXiv:1611.06878v32016Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
Spencer Whitehead, Jacob Phillips, Sean Hendryx
cs.CLcs.CVarXiv:2409.00238v12024Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa +8
cs.LGcs.AIcs.CLarXiv:2502.08640v22025Efficient-CapsNet: Capsule Network with Self-Attention Routing
Vittorio Mazzia, Francesco Salvetti, Marcello Chiaberge
cs.CVcs.AIarXiv:2101.12491v22021Nerfies: Deformable Neural Radiance Fields
Keunhong Park, Utkarsh Sinha, Jonathan T. Barron +4
cs.CVcs.GRarXiv:2011.12948v52020AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
Yunkang Cao, Jiangning Zhang, Luca Frittoli +3
cs.CVarXiv:2407.15795v12024FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image Classification
Zhuo Zheng, Yanfei Zhong, Ailong Ma +1
cs.CVeess.IVarXiv:2011.05670v12020Propagating Confidences through CNNs for Sparse Data Regression
Abdelrahman Eldesokey, Michael Felsberg, Fahad Shahbaz Khan
cs.CVcs.LGarXiv:1805.11913v32018Action Recognition with Image Based CNN Features
Mahdyar Ravanbakhsh, Hossein Mousavi, Mohammad Rastegari +2
cs.CVarXiv:1512.03980v12015DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
Siwei Xia, Li Sun, Tiantian Sun +1
cs.CVarXiv:2505.12427v22025Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos
Zhengze Xu, Mengting Chen, Zhao Wang +6
cs.CVarXiv:2404.17571v120242023 Low-Power Computer Vision Challenge (LPCVC) Summary
Leo Chen, Benjamin Boardley, Ping Hu +27
cs.CVarXiv:2403.07153v12024Orca: The World is in Your Mind
Yihao Wang, Yuheng Ji, Mingyu Cao +54
cs.CVarXiv:2606.30534v32026Is the deconvolution layer the same as a convolutional layer?
Wenzhe Shi, Jose Caballero, Lucas Theis +4
cs.CVarXiv:1609.07009v12016End-to-end Lane Detection through Differentiable Least-Squares Fitting
Wouter Van Gansbeke, Bert De Brabandere, Davy Neven +2
cs.CVarXiv:1902.00293v32019Fruit Quality and Defect Image Classification with Conditional GAN Data Augmentation
Jordan J. Bird, Chloe M. Barnes, Luis J. Manso +2
cs.CVcs.LGeess.IVarXiv:2104.05647v12021Learning from Multimodal and Multitemporal Earth Observation Data for Building Damage Mapping
Bruno Adriano, Naoto Yokoya, Junshi Xia +4
cs.CVcs.LGarXiv:2009.06200v12020From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings
Zahra Anvari, Vassilis Athitsos
cs.CLcs.CVarXiv:2609.17538v12026PRISM: Predictive Representation of Interaction Style and Motion for Social Robot Navigation
Bo-Han Chen, Hiromu Taketsugu, Norimichi Ukita
cs.CVcs.ROarXiv:2609.18125v12026A Multi-Stage model based on YOLOv3 for defect detection in PV panels based on IR and Visible Imaging by Unmanned Aerial Vehicle
Antonio Di Tommaso, Alessandro Betti, Giacomo Fontanelli +1
cs.CVcs.LGarXiv:2111.11709v22021TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning
Yang Liu, Keze Wang, Lingbo Liu +2
cs.CVarXiv:2112.03587v32021In-Context Robot Learning with VLM Agents
Dongzhou Cheng, Taoran Yi, Ye Fang +12
cs.CVcs.ROarXiv:2609.19138v12026FreqMamba: Viewing Mamba from a Frequency Perspective for Image Deraining
Zou Zhen, Yu Hu, Zhao Feng
cs.CVarXiv:2404.09476v22024Energy-Regularized Imitation Learning for Force- and Work-Aware Robotic Manipulation
Toshiki Otani, Hiromu Taketsugu, Norimichi Ukita
cs.ROcs.CVarXiv:2609.18164v12026Learning to Rank for Blind Image Quality Assessment
Fei Gao, Dacheng Tao, Xinbo Gao +1
cs.CVarXiv:1309.0213v32013VOIDD: automatic vessel of intervention dynamic detection in PCI procedures
Ketan Bacchuwar, Jean Cousty, Régis Vaillant +1
cs.CVarXiv:1710.04476v12017MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
Jung Min Lee, Dohyeok Lee, Seokhun Ju +5
cs.ROcs.CVarXiv:2602.03668v32026OminiControl2: Efficient Conditioning for Diffusion Transformers
Zhenxiong Tan, Qiaochu Xue, Xingyi Yang +2
cs.CVcs.AIarXiv:2503.08280v12025Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception
Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns
cs.ROcs.CVcs.LGarXiv:2609.18100v12026FASA: Feature Augmentation and Sampling Adaptation for Long-Tailed Instance Segmentation
Yuhang Zang, Chen Huang, Chen Change Loy
cs.CVarXiv:2102.12867v22021Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds
Vishnu Bindu Balachandran
cs.LGcs.AIcs.CVarXiv:2609.17560v12026MCLC-NET: Multimodal Continual Learning for Leaf Counting
Ruchi Bhatt, Pratibha Kumari, Shreya Bansal +3
cs.CVcs.LGarXiv:2609.18129v12026Temperon: Full-Time SAM Quality at a Third Less Wall-Clock
Stamatis Mastromichalakis
cs.LGcs.CVarXiv:2609.17575v120263D AGSE-VNet: An Automatic Brain Tumor MRI Data Segmentation Framework
Xi Guan, Guang Yang, Jianming Ye +4
cs.AIcs.CVcs.LGarXiv:2107.12046v12021HairCS: Reconstructing Strand-Based Hair from Hair Cards
Zixuan Lu, Tongtong Wang, Yuefan Shen +4
cs.GRcs.CVarXiv:2609.16465v12026Generative Video Motion Editing with 3D Point Tracks
Yao-Chih Lee, Zhoutong Zhang, Jiahui Huang +5
cs.CVarXiv:2512.02015v12025CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework
Yunxiang Fu, Meng Lou, Zicheng Liao +1
cs.LGcs.CVarXiv:2609.17026v12026Different Approaches for Human Activity Recognition: A Survey
Zawar Hussain, Michael Sheng, Wei Emma Zhang
cs.CVarXiv:1906.05074v12019Less is More: Focus Attention for Efficient DETR
Dehua Zheng, Wenhui Dong, Hailin Hu +2
cs.CVcs.AIarXiv:2307.12612v12023A Comprehensive Review of Generative Physical Artificial Intelligence
Satyam Gaba, Krutiksinh Rana, Siva Sai +2
cs.ROcs.AIcs.CLarXiv:2609.18111v12026Data Augmentation for Skin Lesion using Self-Attention based Progressive Generative Adversarial Network
Ibrahim Saad Ali, Mamdouh Farouk Mohamed, Yousef Bassyouni Mahdy
eess.IVcs.CVarXiv:1910.11960v12019Rain rendering for evaluating and improving robustness to bad weather
Maxime Tremblay, Shirsendu Sukanta Halder, Raoul de Charette +1
cs.CVarXiv:2009.03683v12020Multimodal Engagement Analysis from Facial Videos in the Classroom
Ömer Sümer, Patricia Goldberg, Sidney D'Mello +3
cs.CVcs.MMarXiv:2101.04215v22021Bi-Level Routing and Sparse Spatial Attention based Multi-View BEV 3D Object Detection for Autonomous Driving
Jing Zhang, Jiaqi Liu, Zibo Wang
cs.CVcs.AIarXiv:2609.14185v12026Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes
Manglesh Kumar Pandey, Sumit Kumar Banshal
cs.CVcs.LGarXiv:2609.16859v12026FRPSS: Feature Rearrangement in Pre-Shape Space for Single-Image Generation
Yuexing Han, Haoxuan Zhang, Bing Wang
cs.CVarXiv:2609.16594v12026HS-FPN: High Frequency and Spatial Perception FPN for Tiny Object Detection
Zican Shi, Jing Hu, Jie Ren +6
cs.CVarXiv:2412.10116v32024Adaptive AI: Energy Efficient Multi-exit TinyML on Intelligent Vision Systems at the Edge
Luca Crupi, Lorenzo Lamberti, Alessandro Giusti +1
cs.ARcs.CVcs.DCarXiv:2609.11939v12026