Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,501 to 4,560 of 18,817
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
Haomiao Xiong, Zongxin Yang, Jiazuo Yu +4
cs.CVcs.AIarXiv:2501.13468v12025Visual Framing for News Stance Detection via Image Generation
Dahyun Lee, Jiyoung Han, Kunwoo Park
cs.CLcs.AIcs.CVarXiv:2609.00685v12026MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
Xuehui Wang, Zhenyu Wu, JingJing Xie +25
cs.CVcs.CLarXiv:2507.19478v12025OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu +8
cs.CLcs.CVcs.HCarXiv:2410.23218v12024AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
Zhaopeng Gu, Bingke Zhu, Guibo Zhu +3
cs.CVarXiv:2308.15366v42023A Spatio-temporal Transformer for 3D Human Motion Prediction
Emre Aksan, Manuel Kaufmann, Peng Cao +1
cs.CVarXiv:2004.08692v32020Multi-source Distilling Domain Adaptation
Sicheng Zhao, Guangzhi Wang, Shanghang Zhang +7
cs.LGcs.CVstat.MLarXiv:1911.11554v22019ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
David Berthelot, Nicholas Carlini, Ekin D. Cubuk +4
cs.LGcs.CVstat.MLarXiv:1911.09785v22019Multi-Adversarial Domain Adaptation
Zhongyi Pei, Zhangjie Cao, Mingsheng Long +1
cs.CVarXiv:1809.02176v12018Light Field Reconstruction Using Shearlet Transform
Suren Vagharshakyan, Robert Bregovic, Atanas Gotchev
cs.CVarXiv:1509.08969v12015Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Liping Yuan, Jiawei Wang, Haomiao Sun +2
cs.CVcs.AIarXiv:2501.07888v32025Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation
Yuanwang Yang, Buzhen Huang, Zongxuan Ren +2
cs.CVarXiv:2609.00745v12026Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
Ryan Steed, Aylin Caliskan
cs.CYcs.CVarXiv:2010.15052v32020SwiftNet: Real-time Video Object Segmentation
Haochen Wang, Xiaolong Jiang, Haibing Ren +2
cs.CVarXiv:2102.04604v22021Kernel Principal Component Analysis and its Applications in Face Recognition and Active Shape Models
Quan Wang
cs.CVarXiv:1207.3538v32012Arbitrary Shape Scene Text Detection with Adaptive Text Region Representation
Xiaobing Wang, Yingying Jiang, Zhenbo Luo +3
cs.CVarXiv:1905.05980v12019A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
Ahmad Alfan Alfian Irfan, Nur Ahmad Khatim, Mansur Arief
cs.AIcs.CVarXiv:2609.00718v12026Explainable Multi-Loss Distillation Framework for Efficient and Interpretable Shrimp Disease Text Classification
Anh Nguyen Quynh, Khang Nguyen Quoc, Luyl-Da Quach
cs.CVarXiv:2608.29027v12026Visual Personalization Turing Test
Rameen Abdal, James Burgess, Sergey Tulyakov +1
cs.CVarXiv:2601.22680v12026Bolt3D: Generating 3D Scenes in Seconds
Stanislaw Szymanowicz, Jason Y. Zhang, Pratul Srinivasan +6
cs.CVarXiv:2503.14445v22025Geometry-Aware Symmetric Domain Adaptation for Monocular Depth Estimation
Shanshan Zhao, Huan Fu, Mingming Gong +1
cs.CVarXiv:1904.01870v12019A Keypoint-based Global Association Network for Lane Detection
Jinsheng Wang, Yinchao Ma, Shaofei Huang +4
cs.CVarXiv:2204.07335v12022Agent S: An Open Agentic Framework that Uses Computers Like a Human
Saaket Agashe, Jiuzhou Han, Shuyu Gan +3
cs.AIcs.CLcs.CVarXiv:2410.08164v12024A Fusion Approach for Efficient Human Skin Detection
Wei Ren Tan, Chee Seng Chan, Pratheepan Yogarajah +1
cs.CVstat.MLarXiv:1410.3751v12014Vision Foundation Models for Computed Tomography
Suraj Pai, Ibrahim Hadzic, Dennis Bontempi +5
eess.IVcs.CVarXiv:2501.09001v22025Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
Brandon Smart, Chuanxia Zheng, Iro Laina +1
cs.CVcs.LGarXiv:2408.13912v22024Driving on Registers
Ellington Kirby, Alexandre Boulch, Yihong Xu +11
cs.CVcs.AIcs.ROarXiv:2601.05083v22026Improving Sample Quality of Diffusion Models Using Self-Attention Guidance
Susung Hong, Gyuseong Lee, Wooseok Jang +1
cs.CVcs.AIcs.LGarXiv:2210.00939v62022HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
Qi Cai, Jingwen Chen, Chengmin Gao +22
cs.CVcs.MMarXiv:2605.11061v12026Efficient Language-to-Vision Feature Injection for Referring Single-Object Tracking
Han Wang, Yuxuan Liu, Yuhan Sun +5
cs.CVarXiv:2608.29126v12026Semantic Object Parsing with Local-Global Long Short-Term Memory
Xiaodan Liang, Xiaohui Shen, Donglai Xiang +3
cs.CVarXiv:1511.04510v12015WorldEval: World Model as Real-World Robot Policies Evaluator
Yaxuan Li, Yichen Zhu, Junjie Wen +2
cs.ROcs.CVcs.LGarXiv:2505.19017v12025Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
Yuan Yao, Chang Liu, Dezhao Luo +2
cs.CVarXiv:2006.11476v12020Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
Yue Ma, Hongyu Liu, Hongfa Wang +8
cs.CVarXiv:2406.01900v32024Recognizing Fine-Grained and Composite Activities using Hand-Centric Features and Script Data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri +4
cs.CVarXiv:1502.06648v22015Unsupervised Extraction of Video Highlights Via Robust Recurrent Auto-encoders
Huan Yang, Baoyuan Wang, Stephen Lin +3
cs.CVarXiv:1510.01442v12015Test-Time Scaling for Video Diffusion Models via Diagnosis-Guided Candidate Recycling
Hangzhou He, Lunhao Duan, Shanshan Zhao +4
cs.CVarXiv:2608.29322v12026Dynamic Slimmable Network
Changlin Li, Guangrun Wang, Bing Wang +3
cs.CVarXiv:2103.13258v12021Co-Evolutionary Prompt Optimization with Cross-Category Transfer for Zero-Shot Anomaly Detection
Sisi Zhu, Changwei Yu, Renshuai Tao +1
cs.CVarXiv:2608.29467v12026Self-Promoted Prototype Refinement for Few-Shot Class-Incremental Learning
Kai Zhu, Yang Cao, Wei Zhai +2
cs.CVarXiv:2107.08918v12021GENMO: A GENeralist Model for Human MOtion
Jiefeng Li, Jinkun Cao, Haotian Zhang +4
cs.GRcs.AIcs.CVarXiv:2505.01425v12025SGPDFuse: Semantically-Guided Physics-Disentanglement General Multi-Modal Image Fusion
Haozhen Wei, Chengjun Jiang, Yutong Guo +4
cs.CVarXiv:2608.29220v12026Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis
Haoye Dong, Xiaodan Liang, Ke Gong +3
cs.CVarXiv:1810.11610v22018Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation
Liyiming Ke, Xiujun Li, Yonatan Bisk +6
cs.CLcs.CVcs.LGarXiv:1903.02547v22019Foundational feature fusion for conditional flow matching in 6D pose estimation
Amir Hamza, Davide Boscaini, Fabio Poiesi
cs.CVarXiv:2608.29183v12026Language-Conditioned Graph Networks for Relational Reasoning
Ronghang Hu, Anna Rohrbach, Trevor Darrell +1
cs.CVarXiv:1905.04405v22019Anisotropic Convolutional Networks for 3D Semantic Scene Completion
Jie Li, Kai Han, Peng Wang +2
cs.CVarXiv:2004.02122v12020Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
Zeren Jiang, Chuanxia Zheng, Iro Laina +2
cs.CVarXiv:2504.07961v22025Cross-Scale Cost Aggregation for Stereo Matching
Kang Zhang, Yuqiang Fang, Dongbo Min +3
cs.CVarXiv:1403.0316v12014YOLO26: An Analysis of NMS-Free End to End Framework for Real-Time Object Detection
Sudip Chakrabarty
cs.CVcs.AIarXiv:2601.12882v22026See Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval
Xiujun Shu, Wei Wen, Haoqian Wu +5
cs.CVarXiv:2208.08608v22022VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Xilin Wei, Xiaoran Liu, Yuhang Zang +9
cs.CVarXiv:2502.05173v32025A General Two-Step Approach to Learning-Based Hashing
Guosheng Lin, Chunhua Shen, David Suter +1
cs.LGcs.CVarXiv:1309.1853v12013Show, Describe and Conclude: On Exploiting the Structure Information of Chest X-Ray Reports
Baoyu Jing, Zeya Wang, Eric Xing
cs.CLcs.CVeess.IVarXiv:2004.12274v22020Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing
Jian Gao, Chun Gu, Youtian Lin +5
cs.CVcs.GRarXiv:2311.16043v22023Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
Runsen Xu, Weiyao Wang, Hao Tang +5
cs.CVcs.CLarXiv:2505.17015v22025Ovis-U1 Technical Report
Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9
cs.CVcs.AIarXiv:2506.23044v22025SiCloPe: Silhouette-Based Clothed People
Ryota Natsume, Shunsuke Saito, Zeng Huang +4
cs.CVarXiv:1901.00049v22018Reconstructing Personalized Semantic Facial NeRF Models From Monocular Video
Xuan Gao, Chenglai Zhong, Jun Xiang +3
cs.GRcs.CVarXiv:2210.06108v12022InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Xiaoyi Dong, Pan Zhang, Yuhang Zang +21
cs.CVcs.CLarXiv:2404.06512v12024