Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
61 to 120 of 18,795
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
Haoyu Zhao, Zihao Zhang, Jiaxi Gu +10
cs.CVarXiv:2604.09201v12026MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?
Uicheol Jung, Juyoung Hong, Geuntaek Lim +1
cs.CVarXiv:2609.02565v12026Implicit Behavioral Cloning
Pete Florence, Corey Lynch, Andy Zeng +7
cs.ROcs.CVcs.LGarXiv:2109.00137v12021OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
Yiduo Jia, Muzhi Zhu, Hao Zhong +7
cs.CVarXiv:2604.08209v12026ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
Hao Yin, Guangzong Si, Zilei Wang
cs.CVcs.AIarXiv:2503.13107v22025DPA: Decoupling Product-Agnostic Anomaly Representations for Zero-shot Anomaly Generation
Hang Yao, Yansheng Fu, Ming Liu +4
cs.CVarXiv:2609.02075v12026Deep learning cardiac motion analysis for human survival prediction
Ghalib A. Bello, Timothy J. W. Dawes, Jinming Duan +8
cs.LGcs.CVstat.MLarXiv:1810.03382v12018Part-based Graph Convolutional Network for Action Recognition
Kalpit Thakkar, P J Narayanan
cs.CVcs.AIarXiv:1809.04983v12018Breast Tumor Segmentation and Shape Classification in Mammograms using Generative Adversarial and Convolutional Neural Network
Vivek Kumar Singh, Hatem A. Rashwan, Santiago Romani +8
cs.CVarXiv:1809.01687v32018Tensor Methods in Computer Vision and Deep Learning
Yannis Panagakis, Jean Kossaifi, Grigorios G. Chrysos +4
cs.CVarXiv:2107.03436v12021Unified Perceptual Parsing for Scene Understanding
Tete Xiao, Yingcheng Liu, Bolei Zhou +2
cs.CVarXiv:1807.10221v12018From Visual Cues to Spoken Narration: Rethinking Audio Description
Akshita Gupta, Aditya Arora, Federico Tombari +2
cs.CVarXiv:2609.01725v12026Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept Erasure
Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang +2
cs.CVarXiv:2609.01433v12026Synthetic Depth-of-Field with a Single-Camera Mobile Phone
Neal Wadhwa, Rahul Garg, David E. Jacobs +7
cs.CVcs.GRarXiv:1806.04171v12018Summaries:한국어3D Consistent & Robust Segmentation of Cardiac Images by Deep Learning with Spatial Propagation
Qiao Zheng, Hervé Delingette, Nicolas Duchateau +1
cs.CVcs.AIcs.LGarXiv:1804.09400v12018Monocular Depth Estimation from a Single Image: Progress and Opportunities
Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun +4
cs.CVarXiv:2609.01172v12026Automatic Image-Level Morphological Trait Annotation for Organismal Images
Vardaan Pahuja, Samuel Stevens, Alyson East +2
cs.CVcs.AIarXiv:2604.01619v32026Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
Chujie Qin, Zilong Zhang, Zewei Chang +5
cs.CVarXiv:2609.01148v12026Efficient Training of Visual Transformers with Small Datasets
Yahui Liu, Enver Sangineto, Wei Bi +3
cs.CVcs.LGarXiv:2106.03746v22021Endmember-Guided Unmixing Network (EGU-Net): A General Deep Learning Framework for Self-Supervised Hyperspectral Unmixing
Danfeng Hong, Lianru Gao, Jing Yao +4
eess.IVcs.CVarXiv:2105.10194v12021Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
Jingwen Wang, Wenhao Jiang, Lin Ma +2
cs.CVarXiv:1804.00100v22018Revisiting RCNN: On Awakening the Classification Power of Faster RCNN
Bowen Cheng, Yunchao Wei, Honghui Shi +3
cs.CVarXiv:1803.06799v32018ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration
Fengyuan Yang, Luying Huang, Jiazhi Guan +8
cs.CVarXiv:2604.01043v12026LEGO: Learning Edge with Geometry all at Once by Watching Videos
Zhenheng Yang, Peng Wang, Yang Wang +2
cs.CVarXiv:1803.05648v22018TrajectoryMover: Generative Movement of Object Trajectories in Videos
Kiran Chhatre, Hyeonho Jeong, Yulia Gryaditskaya +3
cs.CVarXiv:2603.29092v32026Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli, Soroush Mehraban, Jacob Rommann +3
cs.CVarXiv:2609.00369v12026GEditBench v2: A Human-Aligned Benchmark for General Image Editing
Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7
cs.CVarXiv:2603.28547v12026Visual Interpretability for Deep Learning: a Survey
Quanshi Zhang, Song-Chun Zhu
cs.CVarXiv:1802.00614v22018VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Qijia He, Xunmei Liu, Hammaad Memon +6
cs.CVcs.AIarXiv:2603.24575v22026LaVAN: Localized and Visible Adversarial Noise
Danny Karmon, Daniel Zoran, Yoav Goldberg
cs.CVcs.LGarXiv:1801.02608v22018Maximum Classifier Discrepancy for Unsupervised Domain Adaptation
Kuniaki Saito, Kohei Watanabe, Yoshitaka Ushiku +1
cs.CVarXiv:1712.02560v42017Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model
Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
cs.CVarXiv:2608.29904v12026VITON: An Image-based Virtual Try-on Network
Xintong Han, Zuxuan Wu, Zhe Wu +2
cs.CVarXiv:1711.08447v42017SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
Byungwoo Jeon, Dongyoung Kim, Huiwon Jang +2
cs.CVarXiv:2603.22057v12026Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
Hayeon Kim, Ji Ha Jang, Junghun James Kim +1
cs.CVcs.AIarXiv:2603.22042v32026Sharpness-aware Low dose CT denoising using conditional generative adversarial network
Xin Yi, Paul Babyn
cs.CVarXiv:1708.06453v22017Tinted Frames: Question Framing Blinds Vision-Language Models
Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher +2
cs.CVarXiv:2603.19203v22026A Generic Deep Architecture for Single Image Reflection Removal and Image Smoothing
Qingnan Fan, Jiaolong Yang, Gang Hua +2
cs.CVarXiv:1708.03474v22017Versatile Editing of Video Content, Actions, and Dynamics without Training
Vladimir Kulikov, Roni Paiss, Andrey Voynov +3
cs.CVarXiv:2603.17989v12026Background-Free Objectness Learning for Class-Agnostic Detection
Dania Batool, Liliana Lo Presti, Marco La Cascia +1
cs.CVcs.AIcs.ROarXiv:2608.29232v12026R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection
Yingying Jiang, Xiangyu Zhu, Xiaobing Wang +5
cs.CVarXiv:1706.09579v22017VideoSET: Video Summary Evaluation through Text
Serena Yeung, Alireza Fathi, Li Fei-Fei
cs.CVcs.CLcs.IRarXiv:1406.5824v12014Learning a No-Reference Quality Metric for Single-Image Super-Resolution
Chao Ma, Chih-Yuan Yang, Xiaokang Yang +1
cs.CVarXiv:1612.05890v12016Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
Dennis Menn, Yuedong Yang, Bokun Wang +6
cs.CVarXiv:2603.05811v22026TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
Hanshen Zhu, Yuliang Liu, Xuecheng Wu +7
cs.CVarXiv:2602.20903v32026GENIUS: Generative Fluid Intelligence Evaluation Suite
Ruichuan An, Sihan Yang, Ziyu Guo +8
cs.LGcs.AIcs.CVarXiv:2602.11144v12026AtlasPatch: Efficient Tissue Detection and High-throughput Patch Extraction for Computational Pathology at Scale
Ahmed Alagha, Christopher Leclerc, Yousef Kotp +12
eess.IVcs.CVq-bio.QMarXiv:2602.03998v22026Unified Personalized Reward Model for Vision Generation
Yibin Wang, Yuhang Zang, Feng Han +4
cs.CVarXiv:2602.02380v22026FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia +3
cs.CVarXiv:1612.01925v12016An Empirical Study of World Model Quantization
Zhongqian Fu, Tianyi Zhao, Kai Han +3
cs.LGcs.CVarXiv:2602.02110v12026Understanding Convolutional Neural Networks with A Mathematical Model
C. -C. Jay Kuo
cs.CVarXiv:1609.04112v22016Unsupervised Monocular Depth Estimation with Left-Right Consistency
Clément Godard, Oisin Mac Aodha, Gabriel J. Brostow
cs.CVcs.LGstat.MLarXiv:1609.03677v32016It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces
Leonhard F. Feiner, Manuel Nickel, Martin Menten +6
cs.LGcs.CVarXiv:2608.24518v12026Continual Visual Learning under Evolving Semantic Concept Shift
Ismail Lamaakal, Chaymae Yahyati, Yassine Maleh +2
cs.CVcs.CLarXiv:2608.23903v12026Deep Image Homography Estimation
Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
cs.CVarXiv:1606.03798v12016Improved Techniques for Training GANs
Tim Salimans, Ian Goodfellow, Wojciech Zaremba +3
cs.LGcs.CVcs.NEarXiv:1606.03498v12016LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology
Marie-Lisa Eich, Kai Standvoss, Timo Milbich +30
cs.CVcs.AIcs.LGarXiv:2608.23803v12026Segmentation from Natural Language Expressions
Ronghang Hu, Marcus Rohrbach, Trevor Darrell
cs.CVarXiv:1603.06180v12016Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers
Federico Stella, Fei Jiang, Zhongshi Jiang +4
cs.CVcs.LGarXiv:2608.23410v12026Loopy: Seamless Video Loop Generation via Anchored Looping Shift of Positional Embedding
Haotian Dong, Wenjing Wang, Chen Li +3
cs.CVarXiv:2608.23090v12026