Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
8,341 to 8,400 of 18,866
Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection
Zhongzheng Ren, Zhiding Yu, Xiaodong Yang +4
cs.CVcs.LGeess.IVarXiv:2004.04725v32020DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li +5
cs.CVcs.AIcs.GRarXiv:2409.02095v22024OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection
Jingyang Zhang, Jingkang Yang, Pengyun Wang +9
cs.LGcs.CVarXiv:2306.09301v52023MEOM: Multi-View Expected-OKS Maximization for Human Pose Triangulation
Ziliang Xiong, Henglin Shi, Per-Erik Forssen
cs.CVarXiv:2608.30521v12026A Sanity Check for AI-generated Image Detection
Shilin Yan, Ouxiang Li, Jiayin Cai +4
cs.CVarXiv:2406.19435v320243D Visual Perception for Self-Driving Cars using a Multi-Camera System: Calibration, Mapping, Localization, and Obstacle Detection
Christian Häne, Lionel Heng, Gim Hee Lee +4
cs.CVarXiv:1708.09839v12017FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection
Shaoqing Xu, Dingfu Zhou, Jin Fang +3
cs.CVarXiv:2106.12449v22021Diffusion Probabilistic Models beat GANs on Medical Images
Gustav Müller-Franzes, Jan Moritz Niehues, Firas Khader +8
eess.IVcs.CVarXiv:2212.07501v12022SuperDepth: Self-Supervised, Super-Resolved Monocular Depth Estimation
Sudeep Pillai, Rares Ambrus, Adrien Gaidon
cs.CVcs.AIcs.LGarXiv:1810.01849v12018Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification
Renchun You, Zhiyao Guo, Lei Cui +3
cs.CVarXiv:1912.07872v22019Normalized and Geometry-Aware Self-Attention Network for Image Captioning
Longteng Guo, Jing Liu, Xinxin Zhu +3
cs.CVcs.CLcs.MMarXiv:2003.08897v12020A Style-Aware Content Loss for Real-time HD Style Transfer
Artsiom Sanakoyeu, Dmytro Kotovenko, Sabine Lang +1
cs.CVarXiv:1807.10201v22018Details or Artifacts: A Locally Discriminative Learning Approach to Realistic Image Super-Resolution
Jie Liang, Hui Zeng, Lei Zhang
eess.IVcs.CVarXiv:2203.09195v12022GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction
Yang Chen, Canyu Shen, Xinzhe Rao +5
cs.CVarXiv:2608.29113v12026Apple Flower Detection using Deep Convolutional Networks
Philipe A. Dias, Amy Tabb, Henry Medeiros
cs.CVarXiv:1809.06357v12018Towards Crafting Text Adversarial Samples
Suranjana Samanta, Sameep Mehta
cs.LGcs.AIcs.CLarXiv:1707.02812v12017Deep Edge Guided Recurrent Residual Learning for Image Super-Resolution
Wenhan Yang, Jiashi Feng, Jianchao Yang +4
cs.CVarXiv:1604.08671v22016Attention Correctness in Neural Image Captioning
Chenxi Liu, Junhua Mao, Fei Sha +1
cs.CVcs.CLcs.LGarXiv:1605.09553v22016Person Re-Identification by Discriminative Selection in Video Ranking
Taiqing Wang, Shaogang Gong, Xiatian Zhu +1
cs.CVarXiv:1601.06260v12016Image Restoration using Total Variation Regularized Deep Image Prior
Jiaming Liu, Yu Sun, Xiaojian Xu +1
cs.CVarXiv:1810.12864v12018A Benchmark for Studying Diabetic Retinopathy: Segmentation, Grading, and Transferability
Yi Zhou, Boyang Wang, Lei Huang +2
cs.CVarXiv:2008.09772v32020Training Quantized Nets: A Deeper Understanding
Hao Li, Soham De, Zheng Xu +3
cs.LGcs.CVstat.MLarXiv:1706.02379v32017Unsupervised Generative Adversarial Cross-modal Hashing
Jian Zhang, Yuxin Peng, Mingkuan Yuan
cs.CVarXiv:1712.00358v12017One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic Control
Wenlong Huang, Igor Mordatch, Deepak Pathak
cs.LGcs.CVstat.MLarXiv:2007.04976v12020Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
Jian Ma, Junhao Liang, Chen Chen +1
cs.CVarXiv:2307.11410v22023Knowledge Matters: Radiology Report Generation with General and Specific Knowledge
Shuxin Yang, Xian Wu, Shen Ge +2
eess.IVcs.CLcs.CVarXiv:2112.15009v22021An End-to-End Compression Framework Based on Convolutional Neural Networks
Feng Jiang, Wen Tao, Shaohui Liu +3
cs.CVarXiv:1708.00838v12017LightFuse: Relightable Interactive Gaussian Scene Reconstruction via Multi-Scan Fusion and 2D Gaussian Ray Tracing
Haonan Zhou, Gaoxiang Linghu, Youlin Jia +5
cs.CVarXiv:2608.29269v12026FISICA: A Deployed Service for Plantar-Pressure and Posture Assessment with Ontology-Grounded Recommendation
Juhwan Song, Heejung Kim, Juntae Noh +5
cs.CVcs.HCcs.IRarXiv:2608.29336v12026MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection
Haoyang He, Yuhu Bai, Jiangning Zhang +7
cs.CVarXiv:2404.06564v42024Color Constancy Using CNNs
Simone Bianco, Claudio Cusano, Raimondo Schettini
cs.CVarXiv:1504.04548v12015Diffusion-based Image Translation using Disentangled Style and Content Representation
Gihyun Kwon, Jong Chul Ye
cs.CVcs.AIcs.LGarXiv:2209.15264v22022PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation
Qitao Zhao, Ce Zheng, Mengyuan Liu +2
cs.CVarXiv:2303.17472v12023Min-Entropy Latent Model for Weakly Supervised Object Detection
Fang Wan, Pengxu Wei, Zhenjun Han +2
cs.CVarXiv:1902.06057v12019Semantic Image Synthesis via Diffusion Models
Wengang Zhou, Weilun Wang, Jianmin Bao +4
cs.CVarXiv:2207.00050v42022Explainable Artificial Intelligence (XAI) in Computational Pathology: Definitions, Taxonomy, and Recommendations
Shubham Innani, Suhang You, Adam Shephard +13
cs.AIcs.CVarXiv:2608.28820v12026Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation
Tianrui Hui, Shaofei Huang, Qisong Han +6
cs.CVarXiv:2608.29121v12026Cross-View Image Synthesis using Conditional GANs
Krishna Regmi, Ali Borji
cs.CVarXiv:1803.03396v22018Rigid and Articulated Point Registration with Expectation Conditional Maximization
Radu Horaud, Florence Forbes, Manuel Yguel +2
cs.CVarXiv:2012.05191v12020DCFNet: Discriminant Correlation Filters Network for Visual Tracking
Qiang Wang, Jin Gao, Junliang Xing +2
cs.CVarXiv:1704.04057v12017SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes
Yutao Cui, Chenkai Zeng, Xiaoyu Zhao +3
cs.CVarXiv:2304.05170v22023Automatically Discovering and Learning New Visual Categories with Ranking Statistics
Kai Han, Sylvestre-Alvise Rebuffi, Sebastien Ehrhardt +2
cs.CVarXiv:2002.05714v12020Task-Driven Modular Networks for Zero-Shot Compositional Learning
Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta +1
cs.CVarXiv:1905.05908v12019ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views
Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski +1
cs.CVarXiv:2608.28895v12026Distributed Deep Learning Model for Intelligent Video Surveillance Systems with Edge Computing
Jianguo Chen, Kenli Li, Qingying Deng +2
cs.CVcs.LGarXiv:1904.06400v12019Frequency-Assisted Mamba for Remote Sensing Image Super-Resolution
Yi Xiao, Qiangqiang Yuan, Kui Jiang +3
cs.CVarXiv:2405.04964v22024Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models
Yiting Qu, Xinyue Shen, Xinlei He +3
cs.CVcs.CRcs.CYarXiv:2305.13873v22023Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut
Yangtao Wang, Xi Shen, Shell Hu +3
cs.CVstat.MLarXiv:2202.11539v22022Fully Convolutional Grasp Detection Network with Oriented Anchor Box
Xinwen Zhou, Xuguang Lan, Hanbo Zhang +3
cs.ROcs.CVarXiv:1803.02209v12018Learnable Manifold Alignment (LeMA) : A Semi-supervised Cross-modality Learning Framework for Land Cover and Land Use Classification
Danfeng Hong, Naoto Yokoya, Nan Ge +2
cs.CVarXiv:1901.02838v12019Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing
Yapeng Tian, Dingzeyu Li, Chenliang Xu
cs.CVcs.MMcs.SDarXiv:2007.10558v12020Machine Learning Information Fusion in Earth Observation: A Comprehensive Review of Methods, Applications and Data Sources
S. Salcedo-Sanz, P. Ghamisi, M. Piles +7
cs.CVcs.LGarXiv:2012.05795v12020Vision-Language Models in Remote Sensing: Current Progress and Future Trends
Xiang Li, Congcong Wen, Yuan Hu +2
cs.CVcs.AIarXiv:2305.05726v22023A Perspective Analysis of Handwritten Signature Technology
Moises Diaz, Miguel A. Ferrer, Donato Impedovo +3
cs.CVarXiv:2405.13555v12024MWIR-4-Plastic: The Identification of Complex End-of-Life Industrial Plastic using Mid-wave Infrared Hyperspectral Imaging and Machine Learning
Elias Arbash, Andréa de Lima Ribeiro, Filipa Simões +9
cs.CVcs.LGarXiv:2608.28874v12026Domain Generalization for Medical Imaging Classification with Linear-Dependency Regularization
Haoliang Li, YuFei Wang, Renjie Wan +3
cs.CVcs.LGeess.IVarXiv:2009.12829v32020Video Probabilistic Diffusion Models in Projected Latent Space
Sihyun Yu, Kihyuk Sohn, Subin Kim +1
cs.CVcs.LGarXiv:2302.07685v22023SinNeRF: Training Neural Radiance Fields on Complex Scenes from a Single Image
Dejia Xu, Yifan Jiang, Peihao Wang +3
cs.CVarXiv:2204.00928v22022Deep Optics for Monocular Depth Estimation and 3D Object Detection
Julie Chang, Gordon Wetzstein
cs.CVeess.IVarXiv:1904.08601v12019Omni-frequency Channel-selection Representations for Unsupervised Anomaly Detection
Yufei Liang, Jiangning Zhang, Shiwei Zhao +3
cs.CVarXiv:2203.00259v22022