Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,901 to 3,960 of 18,848
Facial Landmark Detection with Tweaked Convolutional Neural Networks
Yue Wu, Tal Hassner, KangGeon Kim +2
cs.CVarXiv:1511.04031v22015S3CNet: A Sparse Semantic Scene Completion Network for LiDAR Point Clouds
Ran Cheng, Christopher Agia, Yuan Ren +2
cs.CVcs.AIcs.LGarXiv:2012.09242v12020SCAN: Structure Correcting Adversarial Network for Organ Segmentation in Chest X-rays
Wei Dai, Joseph Doyle, Xiaodan Liang +4
cs.CVarXiv:1703.08770v22017EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
Zhili Cheng, Yuge Tu, Ran Li +9
cs.CVcs.CLarXiv:2501.11858v22025AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation
Di You, Fenglin Liu, Shen Ge +3
eess.IVcs.CVarXiv:2203.10095v12022VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
Hongyang Du, Junjie Ye, Xiaoyan Cong +7
cs.CVcs.AIcs.LGarXiv:2601.23286v42026Towards Grand Unification of Object Tracking
Bin Yan, Yi Jiang, Peize Sun +4
cs.CVarXiv:2207.07078v42022Towards Efficient and Scalable Sharpness-Aware Minimization
Yong Liu, Siqi Mai, Xiangning Chen +2
cs.LGcs.AIcs.CVarXiv:2203.02714v12022Template-Based Feature Aggregation Network for Industrial Anomaly Detection
Wei Luo, Haiming Yao, Wenyong Yu
cs.CVarXiv:2603.22874v12026Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching
David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro +3
cs.HCcs.AIcs.CVarXiv:2501.18588v12025Annotation-efficient deep learning for automatic medical image segmentation
Shanshan Wang, Cheng Li, Rongpin Wang +12
eess.IVcs.CVcs.LGarXiv:2012.04885v32020Faster-GS: Analyzing and Improving Gaussian Splatting Optimization
Florian Hahlbohm, Linus Franke, Martin Eisemann +1
cs.CVcs.GRarXiv:2602.09999v12026RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies
Bala Murali Manoghar Sai Sudhakar, Sourab Bapu Sridhar, Sandipan Das +5
cs.ROcs.CVarXiv:2608.28656v12026NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results
Wenbin Zou, Tianyi Liu, Kejun Wu +37
cs.CVarXiv:2604.06945v32026Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Zuyan Liu, Yuhao Dong, Ziwei Liu +3
cs.CVarXiv:2409.12961v42024The Low-Rank Simplicity Bias in Deep Networks
Minyoung Huh, Hossein Mobahi, Richard Zhang +3
cs.LGcs.CVarXiv:2103.10427v42021JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
Shangchen Miao, Ningya Feng, Jialong Wu +4
cs.CVcs.ROarXiv:2602.11832v12026Is Attention Better Than Matrix Decomposition?
Zhengyang Geng, Meng-Hao Guo, Hongxu Chen +3
cs.CVcs.LGarXiv:2109.04553v22021Learning Granularity-Unified Representations for Text-to-Image Person Re-identification
Zhiyin Shao, Xinyu Zhang, Meng Fang +3
cs.CVarXiv:2207.07802v12022Learning to Resize Images for Computer Vision Tasks
Hossein Talebi, Peyman Milanfar
cs.CVcs.LGarXiv:2103.09950v22021UniMate: One Unified Model to Animate Diverse Skeletons
Linzhan Mou, Jiahui Lei, Zhiyang Dou +4
cs.CVcs.GRcs.LGarXiv:2609.05415v12026OrnaStyler: Ornament-Aware Latent Editing for Content-Preserving 3D Stylization
Tomohiro Aizawa, Shigeru Kuriyama, Chunzhi Gu
cs.CVarXiv:2608.29905v12026Learning to Find Eye Region Landmarks for Remote Gaze Estimation in Unconstrained Settings
Seonwook Park, Xucong Zhang, Andreas Bulling +1
cs.CVarXiv:1805.04771v12018Group-Sparse Signal Denoising: Non-Convex Regularization, Convex Optimization
Po-Yu Chen, Ivan W. Selesnick
cs.CVcs.LGstat.MLarXiv:1308.5038v22013Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation
Pietro Morerio, Jacopo Cavazza, Vittorio Murino
cs.CVarXiv:1711.10288v12017SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
Wenchao Sun, Xuewu Lin, Keyu Chen +4
cs.CVarXiv:2603.29163v12026Monocular 3D Human Pose Estimation by Generation and Ordinal Ranking
Saurabh Sharma, Pavan Teja Varigonda, Prashast Bindal +2
cs.CVcs.LGarXiv:1904.01324v22019InverseRenderNet: Learning single image inverse rendering
Ye Yu, William A. P. Smith
cs.CVarXiv:1811.12328v12018OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Tao Zhang, Xiangtai Li, Hao Fei +5
cs.CVarXiv:2406.19389v22024TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
Qianqian Chen, Hyun Bin Kim, Denzel Elden Wijaya +3
cs.CVarXiv:2608.29577v12026Towards Large yet Imperceptible Adversarial Image Perturbations with Perceptual Color Distance
Zhengyu Zhao, Zhuoran Liu, Martha Larson
cs.CVarXiv:1911.02466v22019NepScript Genesis: Neural Architecture Search for Handwritten Devanagari Digit Synthesis
Mausam Gurung, Prabin Neupane, Sajjan Acharya
cs.CVarXiv:2608.29540v12026Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort
cs.LGcs.AIcs.CLarXiv:2306.12929v22023Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
Weijian Ma, Shizhao Sun, Tianyu Yu +3
cs.CVarXiv:2601.01984v12026Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition
Salah Eddine Bekhouche, Hichem Telli, Azeddine Benlamoudi +3
cs.CVarXiv:2603.15818v12026Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing
Bingyan Liu, Chengyu Wang, Tingfeng Cao +2
cs.CVarXiv:2403.03431v12024Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact
Mohsin Bilal, Aadam, Manahil Raza +6
cs.CVarXiv:2502.08333v12025Cascaded deep monocular 3D human pose estimation with evolutionary training data
Shichao Li, Lei Ke, Kevin Pratama +3
cs.CVcs.LGeess.IVarXiv:2006.07778v32020GenAD: Generalized Predictive Model for Autonomous Driving
Jiazhi Yang, Shenyuan Gao, Yihang Qiu +11
cs.CVarXiv:2403.09630v22024MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
Fangyu Liu, Francesco Piccinno, Syrine Krichene +6
cs.CLcs.AIcs.CVarXiv:2212.09662v22022Multi-scale 3D Convolution Network for Video Based Person Re-Identification
Jianing Li, Shiliang Zhang, Tiejun Huang
cs.CVarXiv:1811.07468v12018VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild
Kun Cheng, Xiaodong Cun, Yong Zhang +6
cs.CVarXiv:2211.14758v12022One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6
cs.CVcs.AIarXiv:2609.04190v12026Multi-Scale Representation Learning for Spatial Feature Distributions using Grid Cells
Gengchen Mai, Krzysztof Janowicz, Bo Yan +3
cs.CVcs.AIcs.LGarXiv:2003.00824v12020NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)
Ya-nan Guan, Shaonan Zhang, Hang Guo +55
cs.CVarXiv:2604.11230v12026Reference-Based Sketch Image Colorization using Augmented-Self Reference and Dense Semantic Correspondence
Junsoo Lee, Eungyeup Kim, Yunsung Lee +3
cs.CVarXiv:2005.05207v12020PointAugment: an Auto-Augmentation Framework for Point Cloud Classification
Ruihui Li, Xianzhi Li, Pheng-Ann Heng +1
cs.CVarXiv:2002.10876v22020Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
Fei Shen, Cong Wang, Junyao Gao +4
cs.CVarXiv:2502.09533v12025Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
Xiang Fang, Wanlong Fang, Changshuo Wang
cs.CVcs.AIarXiv:2605.26501v12026LingoQA: Visual Question Answering for Autonomous Driving
Ana-Maria Marcu, Long Chen, Jan Hünermann +9
cs.ROcs.AIcs.CVarXiv:2312.14115v42023Max-Margin Object Detection
Davis E. King
cs.CVarXiv:1502.00046v12015AdaptIS: Adaptive Instance Selection Network
Konstantin Sofiiuk, Olga Barinova, Anton Konushin
cs.CVarXiv:1909.07829v12019RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
Tianyi Wang, Jiazhou Chen, Yiming Xu +7
cs.ROcs.AIcs.CVarXiv:2608.28718v12026ImageBind-LLM: Multi-modality Instruction Tuning
Jiaming Han, Renrui Zhang, Wenqi Shao +14
cs.MMcs.CLcs.CVarXiv:2309.03905v22023Generalized Video Deblurring for Dynamic Scenes
Tae Hyun Kim, Kyoung Mu Lee
cs.CVarXiv:1507.02438v12015MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
Xingyi He, Hao Yu, Sida Peng +4
cs.CVarXiv:2501.07556v12025ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang +21
cs.CVcs.AIcs.CLarXiv:2607.24743v22026Endo-Depth-and-Motion: Reconstruction and Tracking in Endoscopic Videos using Depth Networks and Photometric Constraints
David Recasens, José Lamarca, José M. Fácil +2
cs.CVcs.LGcs.ROarXiv:2103.16525v22021InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
Hao Yu, Haotong Lin, Jiawei Wang +7
cs.CVarXiv:2601.03252v12026Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction
Jiaxin Huang, Yuanbo Yang, Bangbang Yang +3
cs.CVarXiv:2601.04090v22026