Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,901 to 3,960 of 18,848

  1. Facial Landmark Detection with Tweaked Convolutional Neural Networks

    Yue Wu, Tal Hassner, KangGeon Kim +2

    cs.CVarXiv:1511.04031v22015
  2. S3CNet: A Sparse Semantic Scene Completion Network for LiDAR Point Clouds

    Ran Cheng, Christopher Agia, Yuan Ren +2

    cs.CVcs.AIcs.LGarXiv:2012.09242v12020
  3. SCAN: Structure Correcting Adversarial Network for Organ Segmentation in Chest X-rays

    Wei Dai, Joseph Doyle, Xiaodan Liang +4

    cs.CVarXiv:1703.08770v22017
  4. EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

    Zhili Cheng, Yuge Tu, Ran Li +9

    cs.CVcs.CLarXiv:2501.11858v22025
  5. AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation

    Di You, Fenglin Liu, Shen Ge +3

    eess.IVcs.CVarXiv:2203.10095v12022
  6. VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

    Hongyang Du, Junjie Ye, Xiaoyan Cong +7

    cs.CVcs.AIcs.LGarXiv:2601.23286v42026
  7. Towards Grand Unification of Object Tracking

    Bin Yan, Yi Jiang, Peize Sun +4

    cs.CVarXiv:2207.07078v42022
  8. Towards Efficient and Scalable Sharpness-Aware Minimization

    Yong Liu, Siqi Mai, Xiangning Chen +2

    cs.LGcs.AIcs.CVarXiv:2203.02714v12022
  9. Template-Based Feature Aggregation Network for Industrial Anomaly Detection

    Wei Luo, Haiming Yao, Wenyong Yu

    cs.CVarXiv:2603.22874v12026
  10. Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching

    David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro +3

    cs.HCcs.AIcs.CVarXiv:2501.18588v12025
  11. Annotation-efficient deep learning for automatic medical image segmentation

    Shanshan Wang, Cheng Li, Rongpin Wang +12

    eess.IVcs.CVcs.LGarXiv:2012.04885v32020
  12. Faster-GS: Analyzing and Improving Gaussian Splatting Optimization

    Florian Hahlbohm, Linus Franke, Martin Eisemann +1

    cs.CVcs.GRarXiv:2602.09999v12026
  13. RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies

    Bala Murali Manoghar Sai Sudhakar, Sourab Bapu Sridhar, Sandipan Das +5

    cs.ROcs.CVarXiv:2608.28656v12026
  14. NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results

    Wenbin Zou, Tianyi Liu, Kejun Wu +37

    cs.CVarXiv:2604.06945v32026
  15. Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

    Zuyan Liu, Yuhao Dong, Ziwei Liu +3

    cs.CVarXiv:2409.12961v42024
  16. The Low-Rank Simplicity Bias in Deep Networks

    Minyoung Huh, Hossein Mobahi, Richard Zhang +3

    cs.LGcs.CVarXiv:2103.10427v42021
  17. JEPA-VLA: Video Predictive Embedding is Needed for VLA Models

    Shangchen Miao, Ningya Feng, Jialong Wu +4

    cs.CVcs.ROarXiv:2602.11832v12026
  18. Is Attention Better Than Matrix Decomposition?

    Zhengyang Geng, Meng-Hao Guo, Hongxu Chen +3

    cs.CVcs.LGarXiv:2109.04553v22021
  19. Learning Granularity-Unified Representations for Text-to-Image Person Re-identification

    Zhiyin Shao, Xinyu Zhang, Meng Fang +3

    cs.CVarXiv:2207.07802v12022
  20. Learning to Resize Images for Computer Vision Tasks

    Hossein Talebi, Peyman Milanfar

    cs.CVcs.LGarXiv:2103.09950v22021
  21. UniMate: One Unified Model to Animate Diverse Skeletons

    Linzhan Mou, Jiahui Lei, Zhiyang Dou +4

    cs.CVcs.GRcs.LGarXiv:2609.05415v12026
  22. OrnaStyler: Ornament-Aware Latent Editing for Content-Preserving 3D Stylization

    Tomohiro Aizawa, Shigeru Kuriyama, Chunzhi Gu

    cs.CVarXiv:2608.29905v12026
  23. Learning to Find Eye Region Landmarks for Remote Gaze Estimation in Unconstrained Settings

    Seonwook Park, Xucong Zhang, Andreas Bulling +1

    cs.CVarXiv:1805.04771v12018
  24. Group-Sparse Signal Denoising: Non-Convex Regularization, Convex Optimization

    Po-Yu Chen, Ivan W. Selesnick

    cs.CVcs.LGstat.MLarXiv:1308.5038v22013
  25. Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation

    Pietro Morerio, Jacopo Cavazza, Vittorio Murino

    cs.CVarXiv:1711.10288v12017
  26. SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving

    Wenchao Sun, Xuewu Lin, Keyu Chen +4

    cs.CVarXiv:2603.29163v12026
  27. Monocular 3D Human Pose Estimation by Generation and Ordinal Ranking

    Saurabh Sharma, Pavan Teja Varigonda, Prashast Bindal +2

    cs.CVcs.LGarXiv:1904.01324v22019
  28. InverseRenderNet: Learning single image inverse rendering

    Ye Yu, William A. P. Smith

    cs.CVarXiv:1811.12328v12018
  29. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

    Tao Zhang, Xiangtai Li, Hao Fei +5

    cs.CVarXiv:2406.19389v22024
  30. TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection

    Qianqian Chen, Hyun Bin Kim, Denzel Elden Wijaya +3

    cs.CVarXiv:2608.29577v12026
  31. Towards Large yet Imperceptible Adversarial Image Perturbations with Perceptual Color Distance

    Zhengyu Zhao, Zhuoran Liu, Martha Larson

    cs.CVarXiv:1911.02466v22019
  32. NepScript Genesis: Neural Architecture Search for Handwritten Devanagari Digit Synthesis

    Mausam Gurung, Prabin Neupane, Sajjan Acharya

    cs.CVarXiv:2608.29540v12026
  33. Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing

    Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort

    cs.LGcs.AIcs.CLarXiv:2306.12929v22023
  34. Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation

    Weijian Ma, Shizhao Sun, Tianyu Yu +3

    cs.CVarXiv:2601.01984v12026
  35. Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition

    Salah Eddine Bekhouche, Hichem Telli, Azeddine Benlamoudi +3

    cs.CVarXiv:2603.15818v12026
  36. Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

    Bingyan Liu, Chengyu Wang, Tingfeng Cao +2

    cs.CVarXiv:2403.03431v12024
  37. Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact

    Mohsin Bilal, Aadam, Manahil Raza +6

    cs.CVarXiv:2502.08333v12025
  38. Cascaded deep monocular 3D human pose estimation with evolutionary training data

    Shichao Li, Lei Ke, Kevin Pratama +3

    cs.CVcs.LGeess.IVarXiv:2006.07778v32020
  39. GenAD: Generalized Predictive Model for Autonomous Driving

    Jiazhi Yang, Shenyuan Gao, Yihang Qiu +11

    cs.CVarXiv:2403.09630v22024
  40. MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering

    Fangyu Liu, Francesco Piccinno, Syrine Krichene +6

    cs.CLcs.AIcs.CVarXiv:2212.09662v22022
  41. Multi-scale 3D Convolution Network for Video Based Person Re-Identification

    Jianing Li, Shiliang Zhang, Tiejun Huang

    cs.CVarXiv:1811.07468v12018
  42. VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild

    Kun Cheng, Xiaodong Cun, Yong Zhang +6

    cs.CVarXiv:2211.14758v12022
  43. One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

    Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen +6

    cs.CVcs.AIarXiv:2609.04190v12026
  44. Multi-Scale Representation Learning for Spatial Feature Distributions using Grid Cells

    Gengchen Mai, Krzysztof Janowicz, Bo Yan +3

    cs.CVcs.AIcs.LGarXiv:2003.00824v12020
  45. NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)

    Ya-nan Guan, Shaonan Zhang, Hang Guo +55

    cs.CVarXiv:2604.11230v12026
  46. Reference-Based Sketch Image Colorization using Augmented-Self Reference and Dense Semantic Correspondence

    Junsoo Lee, Eungyeup Kim, Yunsung Lee +3

    cs.CVarXiv:2005.05207v12020
  47. PointAugment: an Auto-Augmentation Framework for Point Cloud Classification

    Ruihui Li, Xianzhi Li, Pheng-Ann Heng +1

    cs.CVarXiv:2002.10876v22020
  48. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    Fei Shen, Cong Wang, Junyao Gao +4

    cs.CVarXiv:2502.09533v12025
  49. Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

    Xiang Fang, Wanlong Fang, Changshuo Wang

    cs.CVcs.AIarXiv:2605.26501v12026
  50. LingoQA: Visual Question Answering for Autonomous Driving

    Ana-Maria Marcu, Long Chen, Jan Hünermann +9

    cs.ROcs.AIcs.CVarXiv:2312.14115v42023
  51. Max-Margin Object Detection

    Davis E. King

    cs.CVarXiv:1502.00046v12015
  52. AdaptIS: Adaptive Instance Selection Network

    Konstantin Sofiiuk, Olga Barinova, Anton Konushin

    cs.CVarXiv:1909.07829v12019
  53. RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

    Tianyi Wang, Jiazhou Chen, Yiming Xu +7

    cs.ROcs.AIcs.CVarXiv:2608.28718v12026
  54. ImageBind-LLM: Multi-modality Instruction Tuning

    Jiaming Han, Renrui Zhang, Wenqi Shao +14

    cs.MMcs.CLcs.CVarXiv:2309.03905v22023
  55. Generalized Video Deblurring for Dynamic Scenes

    Tae Hyun Kim, Kyoung Mu Lee

    cs.CVarXiv:1507.02438v12015
  56. MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

    Xingyi He, Hao Yu, Sida Peng +4

    cs.CVarXiv:2501.07556v12025
  57. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    Hangjie Yuan, Yichen Qian, Zhiwei Tang +21

    cs.CVcs.AIcs.CLarXiv:2607.24743v22026
  58. Endo-Depth-and-Motion: Reconstruction and Tracking in Endoscopic Videos using Depth Networks and Photometric Constraints

    David Recasens, José Lamarca, José M. Fácil +2

    cs.CVcs.LGcs.ROarXiv:2103.16525v22021
  59. InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields

    Hao Yu, Haotong Lin, Jiawei Wang +7

    cs.CVarXiv:2601.03252v12026
  60. Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction

    Jiaxin Huang, Yuanbo Yang, Bangbang Yang +3

    cs.CVarXiv:2601.04090v22026