Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,041 to 8,100 of 18,841

  1. Chargrid: Towards Understanding 2D Documents

    Anoop Raveendra Katti, Christian Reisswig, Cordula Guder +4

    cs.CLcs.CVcs.LGarXiv:1809.08799v12018
  2. Detecting events and key actors in multi-person videos

    Vignesh Ramanathan, Jonathan Huang, Sami Abu-El-Haija +3

    cs.CVcs.AIarXiv:1511.02917v22015
  3. Captioning Images Taken by People Who Are Blind

    Danna Gurari, Yinan Zhao, Meng Zhang +1

    cs.CVarXiv:2002.08565v22020
  4. Reconsidering Representation Alignment for Multi-view Clustering

    Daniel J. Trosten, Sigurd Løkse, Robert Jenssen +1

    cs.CVcs.LGarXiv:2103.07738v12021
  5. Co-Separating Sounds of Visual Objects

    Ruohan Gao, Kristen Grauman

    cs.CVcs.MMcs.SDarXiv:1904.07750v22019
  6. Robust machine learning segmentation for large-scale analysis of heterogeneous clinical brain MRI datasets

    Benjamin Billot, Colin Magdamo, You Cheng +3

    eess.IVcs.CVarXiv:2209.02032v22022
  7. Segment Any 3D Gaussians

    Jiazhong Cen, Jiemin Fang, Chen Yang +4

    cs.CVarXiv:2312.00860v32023
  8. MotionSync: Non-Causal Refinement of Causal Tracker for Label-Efficient 3D Perception

    Rahul Ahuja, Bala Murali Manoghar Sai Sudhakar, Shashwata Gupta +3

    cs.CVarXiv:2608.29567v12026
  9. Hierarchical Fine-Grained Image Forgery Detection and Localization

    Xiao Guo, Xiaohong Liu, Zhiyuan Ren +3

    cs.CVarXiv:2303.17111v12023
  10. X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data

    Danfeng Hong, Naoto Yokoya, Gui-Song Xia +2

    cs.CVarXiv:2006.13806v22020
  11. MOON: A Mixed Objective Optimization Network for the Recognition of Facial Attributes

    Ethan Rudd, Manuel Günther, Terrance Boult

    cs.CVarXiv:1603.07027v22016
  12. Local Gradients Smoothing: Defense against localized adversarial attacks

    Muzammal Naseer, Salman H. Khan, Fatih Porikli

    cs.CVarXiv:1807.01216v22018
  13. Multi-camera Realtime 3D Tracking of Multiple Flying Animals

    Andrew D. Straw, Kristin Branson, Titus R. Neumann +1

    cs.CVarXiv:1001.4297v12010
  14. Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction

    Roger Girgis, Florian Golemo, Felipe Codevilla +5

    cs.ROcs.AIcs.CVarXiv:2104.00563v32021
  15. Unsupervised Scale-consistent Depth Learning from Video

    Jia-Wang Bian, Huangying Zhan, Naiyan Wang +5

    cs.CVarXiv:2105.11610v12021
  16. Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance Rendering

    Youngjoong Kwon, Dahun Kim, Duygu Ceylan +1

    cs.CVcs.GRarXiv:2109.07448v12021
  17. Patching open-vocabulary models by interpolating weights

    Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre +5

    cs.CVcs.LGarXiv:2208.05592v22022
  18. Free-form Video Inpainting with 3D Gated Convolution and Temporal PatchGAN

    Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee +1

    cs.CVarXiv:1904.10247v32019
  19. Optimal Transport Aggregation for Visual Place Recognition

    Sergio Izquierdo, Javier Civera

    cs.CVarXiv:2311.15937v22023
  20. H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion

    Hongyi Xu, Thiemo Alldieck, Cristian Sminchisescu

    cs.CVarXiv:2110.13746v22021
  21. DeepSeg: Deep Neural Network Framework for Automatic Brain Tumor Segmentation using Magnetic Resonance FLAIR Images

    Ramy A. Zeineldin, Mohamed E. Karar, Jan Coburger +2

    eess.IVcs.CVarXiv:2004.12333v12020
  22. Unsupervised Domain Adaptation in Semantic Segmentation: a Review

    Marco Toldo, Andrea Maracani, Umberto Michieli +1

    cs.CVcs.LGeess.IVarXiv:2005.10876v12020
  23. From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

    Can Zhang, Baofeng Zhang, Xiaotian Han +5

    cs.CVarXiv:2608.31005v12026
  24. Y-Net: Joint Segmentation and Classification for Diagnosis of Breast Biopsy Images

    Sachin Mehta, Ezgi Mercan, Jamen Bartlett +3

    cs.CVarXiv:1806.01313v12018
  25. PixelIR: Fidelity-Perception Decoupling via Pixel-Space Image-Residual Flow Matching for Efficient One-Step Real-World Super-Resolution

    Bingtian Qiao, Yue Shi, Yong Guo +2

    cs.CVarXiv:2608.30782v12026
  26. Swin Transformer for Fast MRI

    Jiahao Huang, Yingying Fang, Yinzhe Wu +6

    eess.IVcs.AIcs.CVarXiv:2201.03230v22022
  27. DeepEDN: A Deep Learning-based Image Encryption and Decryption Network for Internet of Medical Things

    Yi Ding, Guozheng Wu, Dajiang Chen +4

    cs.CRcs.CVeess.IVarXiv:2004.05523v22020
  28. InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion

    Sirui Xu, Zhengyuan Li, Yu-Xiong Wang +1

    cs.CVcs.AIcs.GRarXiv:2308.16905v12023
  29. Single-Stage 6D Object Pose Estimation

    Yinlin Hu, Pascal Fua, Wei Wang +1

    cs.CVarXiv:1911.08324v22019
  30. Avalanche: an End-to-End Library for Continual Learning

    Vincenzo Lomonaco, Lorenzo Pellegrini, Andrea Cossu +25

    cs.LGcs.AIcs.CVarXiv:2104.00405v12021
  31. Infinite Latent Feature Selection: A Probabilistic Latent Graph-Based Ranking Approach

    Giorgio Roffo, Simone Melzi, Umberto Castellani +1

    cs.CVarXiv:1707.07538v12017
  32. Occlusions, Motion and Depth Boundaries with a Generic Network for Disparity, Optical Flow or Scene Flow Estimation

    Eddy Ilg, Tonmoy Saikia, Margret Keuper +1

    cs.CVarXiv:1808.01838v22018
  33. SegWave: Wavelet-Driven Segmentation of Tampered Regions

    Siddhi Pravin Lipare, Vishesh Kumar, Akshay Agarwal

    cs.CVarXiv:2608.30714v12026
  34. 3D-PRNN: Generating Shape Primitives with Recurrent Neural Networks

    Chuhang Zou, Ersin Yumer, Jimei Yang +2

    cs.CVcs.AIcs.LGarXiv:1708.01648v12017
  35. VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs

    Jingyi He, Sanghwan Kim, Zeynep Akata

    cs.CVarXiv:2608.30705v12026
  36. Failure or Drift? Evaluating Monocular SLAM under Synthetic and Real-World Corruptions

    Abhay Skaria Thomas, Shashank Agnihotri, Margret Keuper

    cs.CVcs.ROarXiv:2608.30690v12026
  37. BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions

    Wenbo Hu, Yifan Xu, Yi Li +3

    cs.CVcs.AIcs.CLarXiv:2308.09936v32023
  38. ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data Scarcity

    Haoran Wang, Matthew Lau, Alec Helbling +7

    cs.CVcs.CRcs.LGarXiv:2608.29510v12026
  39. Packing and Padding: Coupled Multi-index for Accurate Image Retrieval

    Liang Zheng, Shengjin Wang, Ziqiong Liu +1

    cs.CVarXiv:1402.2681v22014
  40. Class Rectification Hard Mining for Imbalanced Deep Learning

    Qi Dong, Shaogang Gong, Xiatian Zhu

    cs.CVarXiv:1712.03162v12017
  41. MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

    Shengbang Tong, David Fan, Jiachen Zhu +7

    cs.CVarXiv:2412.14164v12024
  42. HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned Sampling

    Benjamin Attal, Jia-Bin Huang, Christian Richardt +4

    cs.CVarXiv:2301.02238v22023
  43. Offline Handwritten Signature Verification - Literature Review

    Luiz G. Hafemann, Robert Sabourin, Luiz S. Oliveira

    cs.CVstat.MLarXiv:1507.07909v42015
  44. Open Domain Generalization with Domain-Augmented Meta-Learning

    Yang Shu, Zhangjie Cao, Chenyu Wang +2

    cs.CVcs.LGarXiv:2104.03620v12021
  45. Pruning from Scratch

    Yulong Wang, Xiaolu Zhang, Lingxi Xie +4

    cs.CVarXiv:1909.12579v12019
  46. Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation

    Christian Bailer, Bertram Taetz, Didier Stricker

    cs.CVarXiv:1508.05151v22015
  47. Deep Semantic Face Deblurring

    Ziyi Shen, Wei-Sheng Lai, Tingfa Xu +2

    cs.CVarXiv:1803.03345v22018
  48. Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning

    Nayyer Aafaq, Naveed Akhtar, Wei Liu +2

    cs.CVarXiv:1902.10322v22019
  49. A Self-Supervised Descriptor for Image Copy Detection

    Ed Pizzi, Sreya Dutta Roy, Sugosh Nagavara Ravindra +2

    cs.CVcs.CRcs.LGarXiv:2202.10261v22022
  50. Learning to Describe Differences Between Pairs of Similar Images

    Harsh Jhamtani, Taylor Berg-Kirkpatrick

    cs.CLcs.CVarXiv:1808.10584v12018
  51. Deep Outdoor Illumination Estimation

    Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap +2

    cs.CVarXiv:1611.06403v32016
  52. Adaptive Fusion for RGB-D Salient Object Detection

    Ningning Wang, Xiaojin Gong

    cs.CVarXiv:1901.01369v22019
  53. Contextual Encoder-Decoder Network for Visual Saliency Prediction

    Alexander Kroner, Mario Senden, Kurt Driessens +1

    cs.CVarXiv:1902.06634v42019
  54. Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses

    Christian Rupprecht, Iro Laina, Robert DiPietro +4

    cs.CVarXiv:1612.00197v32016
  55. A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment

    Shuang Ma, Jing Liu, Chang Wen Chen

    cs.CVarXiv:1704.00248v12017
  56. From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real Data

    Ye Liu, Lei Zhu, Shunda Pei +5

    cs.CVarXiv:2108.02934v12021
  57. FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

    Shengyao Chen, Xianbing Sun, Liqing Zhang +1

    cs.CVarXiv:2608.30450v12026
  58. PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

    Youngchae Chee, Hosu Lee, Sungjune Park +2

    cs.CVcs.AIarXiv:2608.30388v12026
  59. Weakly Supervised Video Moment Retrieval From Text Queries

    Niluthpol Chowdhury Mithun, Sujoy Paul, Amit K. Roy-Chowdhury

    cs.CVcs.MMarXiv:1904.03282v22019
  60. You Only Need Adversarial Supervision for Semantic Image Synthesis

    Vadim Sushko, Edgar Schönfeld, Dan Zhang +3

    cs.CVcs.LGeess.IVarXiv:2012.04781v32020