Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

61 to 120 of 18,795

  1. CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

    Haoyu Zhao, Zihao Zhang, Jiaxi Gu +10

    cs.CVarXiv:2604.09201v12026
  2. MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

    Uicheol Jung, Juyoung Hong, Geuntaek Lim +1

    cs.CVarXiv:2609.02565v12026
  3. Implicit Behavioral Cloning

    Pete Florence, Corey Lynch, Andy Zeng +7

    cs.ROcs.CVcs.LGarXiv:2109.00137v12021
  4. OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

    Yiduo Jia, Muzhi Zhu, Hao Zhong +7

    cs.CVarXiv:2604.08209v12026
  5. ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

    Hao Yin, Guangzong Si, Zilei Wang

    cs.CVcs.AIarXiv:2503.13107v22025
  6. DPA: Decoupling Product-Agnostic Anomaly Representations for Zero-shot Anomaly Generation

    Hang Yao, Yansheng Fu, Ming Liu +4

    cs.CVarXiv:2609.02075v12026
  7. Deep learning cardiac motion analysis for human survival prediction

    Ghalib A. Bello, Timothy J. W. Dawes, Jinming Duan +8

    cs.LGcs.CVstat.MLarXiv:1810.03382v12018
  8. Part-based Graph Convolutional Network for Action Recognition

    Kalpit Thakkar, P J Narayanan

    cs.CVcs.AIarXiv:1809.04983v12018
  9. Breast Tumor Segmentation and Shape Classification in Mammograms using Generative Adversarial and Convolutional Neural Network

    Vivek Kumar Singh, Hatem A. Rashwan, Santiago Romani +8

    cs.CVarXiv:1809.01687v32018
  10. Tensor Methods in Computer Vision and Deep Learning

    Yannis Panagakis, Jean Kossaifi, Grigorios G. Chrysos +4

    cs.CVarXiv:2107.03436v12021
  11. Unified Perceptual Parsing for Scene Understanding

    Tete Xiao, Yingcheng Liu, Bolei Zhou +2

    cs.CVarXiv:1807.10221v12018
  12. From Visual Cues to Spoken Narration: Rethinking Audio Description

    Akshita Gupta, Aditya Arora, Federico Tombari +2

    cs.CVarXiv:2609.01725v12026
  13. Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept Erasure

    Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang +2

    cs.CVarXiv:2609.01433v12026
  14. Synthetic Depth-of-Field with a Single-Camera Mobile Phone

    Neal Wadhwa, Rahul Garg, David E. Jacobs +7

    cs.CVcs.GRarXiv:1806.04171v12018
    Summaries:한국어
  15. 3D Consistent & Robust Segmentation of Cardiac Images by Deep Learning with Spatial Propagation

    Qiao Zheng, Hervé Delingette, Nicolas Duchateau +1

    cs.CVcs.AIcs.LGarXiv:1804.09400v12018
  16. Monocular Depth Estimation from a Single Image: Progress and Opportunities

    Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun +4

    cs.CVarXiv:2609.01172v12026
  17. Automatic Image-Level Morphological Trait Annotation for Organismal Images

    Vardaan Pahuja, Samuel Stevens, Alyson East +2

    cs.CVcs.AIarXiv:2604.01619v32026
  18. Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement

    Chujie Qin, Zilong Zhang, Zewei Chang +5

    cs.CVarXiv:2609.01148v12026
  19. Efficient Training of Visual Transformers with Small Datasets

    Yahui Liu, Enver Sangineto, Wei Bi +3

    cs.CVcs.LGarXiv:2106.03746v22021
  20. Endmember-Guided Unmixing Network (EGU-Net): A General Deep Learning Framework for Self-Supervised Hyperspectral Unmixing

    Danfeng Hong, Lianru Gao, Jing Yao +4

    eess.IVcs.CVarXiv:2105.10194v12021
  21. Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning

    Jingwen Wang, Wenhao Jiang, Lin Ma +2

    cs.CVarXiv:1804.00100v22018
  22. Revisiting RCNN: On Awakening the Classification Power of Faster RCNN

    Bowen Cheng, Yunchao Wei, Honghui Shi +3

    cs.CVarXiv:1803.06799v32018
  23. ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

    Fengyuan Yang, Luying Huang, Jiazhi Guan +8

    cs.CVarXiv:2604.01043v12026
  24. LEGO: Learning Edge with Geometry all at Once by Watching Videos

    Zhenheng Yang, Peng Wang, Yang Wang +2

    cs.CVarXiv:1803.05648v22018
  25. TrajectoryMover: Generative Movement of Object Trajectories in Videos

    Kiran Chhatre, Hyeonho Jeong, Yulia Gryaditskaya +3

    cs.CVarXiv:2603.29092v32026
  26. Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

    Vida Adeli, Soroush Mehraban, Jacob Rommann +3

    cs.CVarXiv:2609.00369v12026
  27. GEditBench v2: A Human-Aligned Benchmark for General Image Editing

    Zhangqi Jiang, Zheng Sun, Xianfang Zeng +7

    cs.CVarXiv:2603.28547v12026
  28. Visual Interpretability for Deep Learning: a Survey

    Quanshi Zhang, Song-Chun Zhu

    cs.CVarXiv:1802.00614v22018
  29. VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models

    Qijia He, Xunmei Liu, Hammaad Memon +6

    cs.CVcs.AIarXiv:2603.24575v22026
  30. LaVAN: Localized and Visible Adversarial Noise

    Danny Karmon, Daniel Zoran, Yoav Goldberg

    cs.CVcs.LGarXiv:1801.02608v22018
  31. Maximum Classifier Discrepancy for Unsupervised Domain Adaptation

    Kuniaki Saito, Kohei Watanabe, Yoshitaka Ushiku +1

    cs.CVarXiv:1712.02560v42017
  32. Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model

    Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh

    cs.CVarXiv:2608.29904v12026
  33. VITON: An Image-based Virtual Try-on Network

    Xintong Han, Zuxuan Wu, Zhe Wu +2

    cs.CVarXiv:1711.08447v42017
  34. SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

    Byungwoo Jeon, Dongyoung Kim, Huiwon Jang +2

    cs.CVarXiv:2603.22057v12026
  35. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

    Hayeon Kim, Ji Ha Jang, Junghun James Kim +1

    cs.CVcs.AIarXiv:2603.22042v32026
  36. Sharpness-aware Low dose CT denoising using conditional generative adversarial network

    Xin Yi, Paul Babyn

    cs.CVarXiv:1708.06453v22017
  37. Tinted Frames: Question Framing Blinds Vision-Language Models

    Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher +2

    cs.CVarXiv:2603.19203v22026
  38. A Generic Deep Architecture for Single Image Reflection Removal and Image Smoothing

    Qingnan Fan, Jiaolong Yang, Gang Hua +2

    cs.CVarXiv:1708.03474v22017
  39. Versatile Editing of Video Content, Actions, and Dynamics without Training

    Vladimir Kulikov, Roni Paiss, Andrey Voynov +3

    cs.CVarXiv:2603.17989v12026
  40. Background-Free Objectness Learning for Class-Agnostic Detection

    Dania Batool, Liliana Lo Presti, Marco La Cascia +1

    cs.CVcs.AIcs.ROarXiv:2608.29232v12026
  41. R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection

    Yingying Jiang, Xiangyu Zhu, Xiaobing Wang +5

    cs.CVarXiv:1706.09579v22017
  42. VideoSET: Video Summary Evaluation through Text

    Serena Yeung, Alireza Fathi, Li Fei-Fei

    cs.CVcs.CLcs.IRarXiv:1406.5824v12014
  43. Learning a No-Reference Quality Metric for Single-Image Super-Resolution

    Chao Ma, Chih-Yuan Yang, Xiaokang Yang +1

    cs.CVarXiv:1612.05890v12016
  44. Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery

    Dennis Menn, Yuedong Yang, Bokun Wang +6

    cs.CVarXiv:2603.05811v22026
  45. TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

    Hanshen Zhu, Yuliang Liu, Xuecheng Wu +7

    cs.CVarXiv:2602.20903v32026
  46. GENIUS: Generative Fluid Intelligence Evaluation Suite

    Ruichuan An, Sihan Yang, Ziyu Guo +8

    cs.LGcs.AIcs.CVarXiv:2602.11144v12026
  47. AtlasPatch: Efficient Tissue Detection and High-throughput Patch Extraction for Computational Pathology at Scale

    Ahmed Alagha, Christopher Leclerc, Yousef Kotp +12

    eess.IVcs.CVq-bio.QMarXiv:2602.03998v22026
  48. Unified Personalized Reward Model for Vision Generation

    Yibin Wang, Yuhang Zang, Feng Han +4

    cs.CVarXiv:2602.02380v22026
  49. FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks

    Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia +3

    cs.CVarXiv:1612.01925v12016
  50. An Empirical Study of World Model Quantization

    Zhongqian Fu, Tianyi Zhao, Kai Han +3

    cs.LGcs.CVarXiv:2602.02110v12026
  51. Understanding Convolutional Neural Networks with A Mathematical Model

    C. -C. Jay Kuo

    cs.CVarXiv:1609.04112v22016
  52. Unsupervised Monocular Depth Estimation with Left-Right Consistency

    Clément Godard, Oisin Mac Aodha, Gabriel J. Brostow

    cs.CVcs.LGstat.MLarXiv:1609.03677v32016
  53. It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces

    Leonhard F. Feiner, Manuel Nickel, Martin Menten +6

    cs.LGcs.CVarXiv:2608.24518v12026
  54. Continual Visual Learning under Evolving Semantic Concept Shift

    Ismail Lamaakal, Chaymae Yahyati, Yassine Maleh +2

    cs.CVcs.CLarXiv:2608.23903v12026
  55. Deep Image Homography Estimation

    Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich

    cs.CVarXiv:1606.03798v12016
  56. Improved Techniques for Training GANs

    Tim Salimans, Ian Goodfellow, Wojciech Zaremba +3

    cs.LGcs.CVcs.NEarXiv:1606.03498v12016
  57. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

    Marie-Lisa Eich, Kai Standvoss, Timo Milbich +30

    cs.CVcs.AIcs.LGarXiv:2608.23803v12026
  58. Segmentation from Natural Language Expressions

    Ronghang Hu, Marcus Rohrbach, Trevor Darrell

    cs.CVarXiv:1603.06180v12016
  59. Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

    Federico Stella, Fei Jiang, Zhongshi Jiang +4

    cs.CVcs.LGarXiv:2608.23410v12026
  60. Loopy: Seamless Video Loop Generation via Anchored Looping Shift of Positional Embedding

    Haotian Dong, Wenjing Wang, Chen Li +3

    cs.CVarXiv:2608.23090v12026