Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1 to 60 of 18,796
PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models
Mingde Yao, Zhiyuan You, King-Man Tam +2
cs.CVarXiv:2602.22809v32026VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
Xiaoyan Cong, Haotian Yang, Angtian Wang +4
cs.CVarXiv:2512.16906v12025MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
Yuhuan Yang, Chaofan Ma, Zhenjie Mao +3
cs.CVarXiv:2506.23283v12025Agentic Knowledgeable Self-awareness
Shuofei Qiao, Zhisong Qiu, Baochang Ren +8
cs.CLcs.AIcs.CVarXiv:2504.03553v22025MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation
Chen Liu, Yasutaka Furukawa
cs.CVarXiv:1902.04478v12019AdaCoSeg: Adaptive Shape Co-Segmentation with Group Consistency Loss
Chenyang Zhu, Kai Xu, Siddhartha Chaudhuri +3
cs.CVcs.GRarXiv:1903.10297v52019PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing
Hehe Fan, Yi Yang
cs.CVarXiv:1910.08287v22019LSANet: Feature Learning on Point Sets by Local Spatial Aware Layer
Lin-Zhuo Chen, Xuan-Yi Li, Deng-Ping Fan +3
cs.CVarXiv:1905.05442v32019Intern-S1: A Scientific Multimodal Foundation Model
Lei Bai, Zhongrui Cai, Yuhang Cao +174
cs.LGcs.CLcs.CVarXiv:2508.15763v22025Rehearsal-Free Continual Learning over Small Non-I.I.D. Batches
Vincenzo Lomonaco, Davide Maltoni, Lorenzo Pellegrini
cs.LGcs.CVcs.NEarXiv:1907.03799v32019JAFAR: Jack up Any Feature at Any Resolution
Paul Couairon, Loick Chambon, Louis Serrano +3
cs.CVeess.IVarXiv:2506.11136v32025GRAPE: Generalizing Robot Policy via Preference Alignment
Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen +7
cs.ROcs.CVcs.LGarXiv:2411.19309v22024Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
Davide Caffagni, Sara Sarto, Marcella Cornia +5
cs.CVcs.AIcs.CLarXiv:2512.15885v12025Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
Zihao Yue, Liang Zhang, Qin Jin
cs.CLcs.CVarXiv:2402.14545v22024Generating Images Part by Part with Composite Generative Adversarial Networks
Hanock Kwak, Byoung-Tak Zhang
cs.AIcs.CVcs.LGarXiv:1607.05387v22016Deep Learning for Omnidirectional Vision: A Survey and New Perspectives
Hao Ai, Zidong Cao, Jinjing Zhu +3
cs.CVarXiv:2205.10468v22022RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
Maëlic Neau
cs.CVarXiv:2609.12552v12026NearID: Identity Representation Learning via Near-identity Distractors
Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey +2
cs.CVarXiv:2604.01973v22026VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
Jiarong Liang, Max Ku, Ka-Hei Hui +2
cs.CVcs.AIarXiv:2602.13294v32026AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
Xintong Zhang, Xiaowen Zhang, Jingrong Wu +8
cs.CVarXiv:2602.02676v32026Domain Generalization by Mutual-Information Regularization with Pre-trained Models
Junbum Cha, Kyungjae Lee, Sungrae Park +1
cs.LGcs.CVarXiv:2203.10789v22022Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection
Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang +2
cs.CVarXiv:2203.02688v12022Visual Prompt Tuning
Menglin Jia, Luming Tang, Bor-Chun Chen +4
cs.CVarXiv:2203.12119v22022Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at Scale
Ram Ramrakhya, Eric Undersander, Dhruv Batra +1
cs.AIcs.CVcs.ROarXiv:2204.03514v22022Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion
Xiang Chen, Ningyu Zhang, Lei Li +6
cs.CLcs.AIcs.CVarXiv:2205.02357v52022MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
Zhenyu Pan, Han Liu
cs.CVcs.AIarXiv:2503.18470v22025Comprehending and Ordering Semantics for Image Captioning
Yehao Li, Yingwei Pan, Ting Yao +1
cs.CVcs.AIcs.CLarXiv:2206.06930v12022CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation
Zhihao Li, Jianzhuang Liu, Zhensong Zhang +2
cs.CVarXiv:2208.00571v22022Curriculum Temperature for Knowledge Distillation
Zheng Li, Xiang Li, Lingfeng Yang +5
cs.CVarXiv:2211.16231v32022One-Stage Cascade Refinement Networks for Infrared Small Target Detection
Yimian Dai, Xiang Li, Fei Zhou +3
cs.CVarXiv:2212.08472v22022Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Yang Chen, Hexiang Hu, Yi Luan +4
cs.CVcs.AIcs.CLarXiv:2302.11713v52023LayoutDM: Discrete Diffusion Model for Controllable Layout Generation
Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra +2
cs.CVcs.GRarXiv:2303.08137v12023FreeDoM: Training-Free Energy-Guided Conditional Diffusion Model
Jiwen Yu, Yinhuai Wang, Chen Zhao +2
cs.CVarXiv:2303.09833v12023LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Renrui Zhang, Jiaming Han, Chris Liu +7
cs.CVcs.AIcs.CLarXiv:2303.16199v32023LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Peng Gao, Jiaming Han, Renrui Zhang +9
cs.CVcs.AIcs.CLarXiv:2304.15010v12023PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao +4
cs.CVarXiv:2305.10415v62023Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark
Shuyu Yang, Yinan Zhou, Yaxiong Wang +3
cs.CVcs.MMarXiv:2306.02898v42023RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
Zilun Zhang, Tiancheng Zhao, Yulong Guo +1
cs.CVcs.AIcs.CLarXiv:2306.11300v52023GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization
Vicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak Shah
cs.CVcs.LGarXiv:2309.16020v22023Gibbs Sampling with People
Peter M. C. Harrison, Raja Marjieh, Federico Adolfi +5
q-bio.NCcs.AIcs.CVarXiv:2008.02595v22020Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
Rohit Girdhar, Mannat Singh, Andrew Brown +7
cs.CVcs.AIcs.GRarXiv:2311.10709v22023HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Qifan Yu, Juncheng Li, Longhui Wei +6
cs.CVcs.AIarXiv:2311.13614v22023VBench: Comprehensive Benchmark Suite for Video Generative Models
Ziqi Huang, Yinan He, Jiashuo Yu +13
cs.CVarXiv:2311.17982v12023i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu +4
cs.CVarXiv:2606.11289v12026Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
Yichen Feng, Yuetai Li, Chunjiang Liu +14
cs.CVcs.AIcs.HCarXiv:2605.12684v12026A Consistent and Efficient Evaluation Strategy for Attribution Methods
Yao Rong, Tobias Leemann, Vadim Borisov +2
cs.CVcs.LGarXiv:2202.00449v22022HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Xiwu Chen +4
cs.CVarXiv:2604.28196v12026Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
Sudong Wang, Weiquan Huang, Xiaomin Yu +9
cs.CVcs.AIcs.CLarXiv:2604.28123v32026BS-Nets: An End-to-End Framework For Band Selection of Hyperspectral Image
Yaoming Cai, Xiaobo Liu, Zhihua Cai
cs.CVcs.LGarXiv:1904.08269v12019A Closer Look at Few-shot Classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira +2
cs.CVarXiv:1904.04232v22019Knowledge-driven Encode, Retrieve, Paraphrase for Medical Image Report Generation
Christy Y. Li, Xiaodan Liang, Zhiting Hu +1
cs.CVarXiv:1903.10122v12019Expansion-Squeeze-Excitation Fusion Network for Elderly Activity Recognition
Xiangbo Shu, Jiawen Yang, Rui Yan +1
cs.CVstat.MLarXiv:2112.10992v22021Class-incremental Learning via Deep Model Consolidation
Junting Zhang, Jie Zhang, Shalini Ghosh +5
cs.CVcs.LGarXiv:1903.07864v42019ICON: Implicit Clothed humans Obtained from Normals
Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas +1
cs.CVcs.AIcs.GRarXiv:2112.09127v22021DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction
Shiyan Su, Ruyi Zha, Danli Shi +2
eess.IVcs.CVarXiv:2604.21518v12026Extracting Triangular 3D Models, Materials, and Lighting From Images
Jacob Munkberg, Jon Hasselgren, Tianchang Shen +5
cs.CVcs.GRarXiv:2111.12503v52021Exploiting Unlabeled Data in CNNs by Self-supervised Learning to Rank
Xialei Liu, Joost van de Weijer, Andrew D. Bagdanov
cs.CVarXiv:1902.06285v12019VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
Jiangang Zhu, Zheng Wang, Bin Zhu +2
cs.CVcs.AIarXiv:2609.04948v12026ABO: Dataset and Benchmarks for Real-World 3D Object Understanding
Jasmine Collins, Shubham Goel, Kenan Deng +9
cs.CVcs.AIcs.GRarXiv:2110.06199v22021PAPT++: Risk-Aware Adversarial Tuning and Generation for Single Domain Generalization
Zhipeng Xu, De Cheng, Xinyang Jiang +5
cs.CVarXiv:2609.04837v12026