Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,041 to 8,100 of 18,841
Chargrid: Towards Understanding 2D Documents
Anoop Raveendra Katti, Christian Reisswig, Cordula Guder +4
cs.CLcs.CVcs.LGarXiv:1809.08799v12018Detecting events and key actors in multi-person videos
Vignesh Ramanathan, Jonathan Huang, Sami Abu-El-Haija +3
cs.CVcs.AIarXiv:1511.02917v22015Captioning Images Taken by People Who Are Blind
Danna Gurari, Yinan Zhao, Meng Zhang +1
cs.CVarXiv:2002.08565v22020Reconsidering Representation Alignment for Multi-view Clustering
Daniel J. Trosten, Sigurd Løkse, Robert Jenssen +1
cs.CVcs.LGarXiv:2103.07738v12021Co-Separating Sounds of Visual Objects
Ruohan Gao, Kristen Grauman
cs.CVcs.MMcs.SDarXiv:1904.07750v22019Robust machine learning segmentation for large-scale analysis of heterogeneous clinical brain MRI datasets
Benjamin Billot, Colin Magdamo, You Cheng +3
eess.IVcs.CVarXiv:2209.02032v22022Segment Any 3D Gaussians
Jiazhong Cen, Jiemin Fang, Chen Yang +4
cs.CVarXiv:2312.00860v32023MotionSync: Non-Causal Refinement of Causal Tracker for Label-Efficient 3D Perception
Rahul Ahuja, Bala Murali Manoghar Sai Sudhakar, Shashwata Gupta +3
cs.CVarXiv:2608.29567v12026Hierarchical Fine-Grained Image Forgery Detection and Localization
Xiao Guo, Xiaohong Liu, Zhiyuan Ren +3
cs.CVarXiv:2303.17111v12023X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data
Danfeng Hong, Naoto Yokoya, Gui-Song Xia +2
cs.CVarXiv:2006.13806v22020MOON: A Mixed Objective Optimization Network for the Recognition of Facial Attributes
Ethan Rudd, Manuel Günther, Terrance Boult
cs.CVarXiv:1603.07027v22016Local Gradients Smoothing: Defense against localized adversarial attacks
Muzammal Naseer, Salman H. Khan, Fatih Porikli
cs.CVarXiv:1807.01216v22018Multi-camera Realtime 3D Tracking of Multiple Flying Animals
Andrew D. Straw, Kristin Branson, Titus R. Neumann +1
cs.CVarXiv:1001.4297v12010Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction
Roger Girgis, Florian Golemo, Felipe Codevilla +5
cs.ROcs.AIcs.CVarXiv:2104.00563v32021Unsupervised Scale-consistent Depth Learning from Video
Jia-Wang Bian, Huangying Zhan, Naiyan Wang +5
cs.CVarXiv:2105.11610v12021Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance Rendering
Youngjoong Kwon, Dahun Kim, Duygu Ceylan +1
cs.CVcs.GRarXiv:2109.07448v12021Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre +5
cs.CVcs.LGarXiv:2208.05592v22022Free-form Video Inpainting with 3D Gated Convolution and Temporal PatchGAN
Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee +1
cs.CVarXiv:1904.10247v32019Optimal Transport Aggregation for Visual Place Recognition
Sergio Izquierdo, Javier Civera
cs.CVarXiv:2311.15937v22023H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion
Hongyi Xu, Thiemo Alldieck, Cristian Sminchisescu
cs.CVarXiv:2110.13746v22021DeepSeg: Deep Neural Network Framework for Automatic Brain Tumor Segmentation using Magnetic Resonance FLAIR Images
Ramy A. Zeineldin, Mohamed E. Karar, Jan Coburger +2
eess.IVcs.CVarXiv:2004.12333v12020Unsupervised Domain Adaptation in Semantic Segmentation: a Review
Marco Toldo, Andrea Maracani, Umberto Michieli +1
cs.CVcs.LGeess.IVarXiv:2005.10876v12020From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents
Can Zhang, Baofeng Zhang, Xiaotian Han +5
cs.CVarXiv:2608.31005v12026Y-Net: Joint Segmentation and Classification for Diagnosis of Breast Biopsy Images
Sachin Mehta, Ezgi Mercan, Jamen Bartlett +3
cs.CVarXiv:1806.01313v12018PixelIR: Fidelity-Perception Decoupling via Pixel-Space Image-Residual Flow Matching for Efficient One-Step Real-World Super-Resolution
Bingtian Qiao, Yue Shi, Yong Guo +2
cs.CVarXiv:2608.30782v12026Swin Transformer for Fast MRI
Jiahao Huang, Yingying Fang, Yinzhe Wu +6
eess.IVcs.AIcs.CVarXiv:2201.03230v22022DeepEDN: A Deep Learning-based Image Encryption and Decryption Network for Internet of Medical Things
Yi Ding, Guozheng Wu, Dajiang Chen +4
cs.CRcs.CVeess.IVarXiv:2004.05523v22020InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang +1
cs.CVcs.AIcs.GRarXiv:2308.16905v12023Single-Stage 6D Object Pose Estimation
Yinlin Hu, Pascal Fua, Wei Wang +1
cs.CVarXiv:1911.08324v22019Avalanche: an End-to-End Library for Continual Learning
Vincenzo Lomonaco, Lorenzo Pellegrini, Andrea Cossu +25
cs.LGcs.AIcs.CVarXiv:2104.00405v12021Infinite Latent Feature Selection: A Probabilistic Latent Graph-Based Ranking Approach
Giorgio Roffo, Simone Melzi, Umberto Castellani +1
cs.CVarXiv:1707.07538v12017Occlusions, Motion and Depth Boundaries with a Generic Network for Disparity, Optical Flow or Scene Flow Estimation
Eddy Ilg, Tonmoy Saikia, Margret Keuper +1
cs.CVarXiv:1808.01838v22018SegWave: Wavelet-Driven Segmentation of Tampered Regions
Siddhi Pravin Lipare, Vishesh Kumar, Akshay Agarwal
cs.CVarXiv:2608.30714v120263D-PRNN: Generating Shape Primitives with Recurrent Neural Networks
Chuhang Zou, Ersin Yumer, Jimei Yang +2
cs.CVcs.AIcs.LGarXiv:1708.01648v12017VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs
Jingyi He, Sanghwan Kim, Zeynep Akata
cs.CVarXiv:2608.30705v12026Failure or Drift? Evaluating Monocular SLAM under Synthetic and Real-World Corruptions
Abhay Skaria Thomas, Shashank Agnihotri, Margret Keuper
cs.CVcs.ROarXiv:2608.30690v12026BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
Wenbo Hu, Yifan Xu, Yi Li +3
cs.CVcs.AIcs.CLarXiv:2308.09936v32023ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data Scarcity
Haoran Wang, Matthew Lau, Alec Helbling +7
cs.CVcs.CRcs.LGarXiv:2608.29510v12026Packing and Padding: Coupled Multi-index for Accurate Image Retrieval
Liang Zheng, Shengjin Wang, Ziqiong Liu +1
cs.CVarXiv:1402.2681v22014Class Rectification Hard Mining for Imbalanced Deep Learning
Qi Dong, Shaogang Gong, Xiatian Zhu
cs.CVarXiv:1712.03162v12017MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Shengbang Tong, David Fan, Jiachen Zhu +7
cs.CVarXiv:2412.14164v12024HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned Sampling
Benjamin Attal, Jia-Bin Huang, Christian Richardt +4
cs.CVarXiv:2301.02238v22023Offline Handwritten Signature Verification - Literature Review
Luiz G. Hafemann, Robert Sabourin, Luiz S. Oliveira
cs.CVstat.MLarXiv:1507.07909v42015Open Domain Generalization with Domain-Augmented Meta-Learning
Yang Shu, Zhangjie Cao, Chenyu Wang +2
cs.CVcs.LGarXiv:2104.03620v12021Pruning from Scratch
Yulong Wang, Xiaolu Zhang, Lingxi Xie +4
cs.CVarXiv:1909.12579v12019Flow Fields: Dense Correspondence Fields for Highly Accurate Large Displacement Optical Flow Estimation
Christian Bailer, Bertram Taetz, Didier Stricker
cs.CVarXiv:1508.05151v22015Deep Semantic Face Deblurring
Ziyi Shen, Wei-Sheng Lai, Tingfa Xu +2
cs.CVarXiv:1803.03345v22018Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning
Nayyer Aafaq, Naveed Akhtar, Wei Liu +2
cs.CVarXiv:1902.10322v22019A Self-Supervised Descriptor for Image Copy Detection
Ed Pizzi, Sreya Dutta Roy, Sugosh Nagavara Ravindra +2
cs.CVcs.CRcs.LGarXiv:2202.10261v22022Learning to Describe Differences Between Pairs of Similar Images
Harsh Jhamtani, Taylor Berg-Kirkpatrick
cs.CLcs.CVarXiv:1808.10584v12018Deep Outdoor Illumination Estimation
Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap +2
cs.CVarXiv:1611.06403v32016Adaptive Fusion for RGB-D Salient Object Detection
Ningning Wang, Xiaojin Gong
cs.CVarXiv:1901.01369v22019Contextual Encoder-Decoder Network for Visual Saliency Prediction
Alexander Kroner, Mario Senden, Kurt Driessens +1
cs.CVarXiv:1902.06634v42019Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
Christian Rupprecht, Iro Laina, Robert DiPietro +4
cs.CVarXiv:1612.00197v32016A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment
Shuang Ma, Jing Liu, Chang Wen Chen
cs.CVarXiv:1704.00248v12017From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real Data
Ye Liu, Lei Zhu, Shunda Pei +5
cs.CVarXiv:2108.02934v12021FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On
Shengyao Chen, Xianbing Sun, Liqing Zhang +1
cs.CVarXiv:2608.30450v12026PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning
Youngchae Chee, Hosu Lee, Sungjune Park +2
cs.CVcs.AIarXiv:2608.30388v12026Weakly Supervised Video Moment Retrieval From Text Queries
Niluthpol Chowdhury Mithun, Sujoy Paul, Amit K. Roy-Chowdhury
cs.CVcs.MMarXiv:1904.03282v22019You Only Need Adversarial Supervision for Semantic Image Synthesis
Vadim Sushko, Edgar Schönfeld, Dan Zhang +3
cs.CVcs.LGeess.IVarXiv:2012.04781v32020