Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,341 to 8,400 of 18,866

  1. Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection

    Zhongzheng Ren, Zhiding Yu, Xiaodong Yang +4

    cs.CVcs.LGeess.IVarXiv:2004.04725v32020
  2. DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

    Wenbo Hu, Xiangjun Gao, Xiaoyu Li +5

    cs.CVcs.AIcs.GRarXiv:2409.02095v22024
  3. OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection

    Jingyang Zhang, Jingkang Yang, Pengyun Wang +9

    cs.LGcs.CVarXiv:2306.09301v52023
  4. MEOM: Multi-View Expected-OKS Maximization for Human Pose Triangulation

    Ziliang Xiong, Henglin Shi, Per-Erik Forssen

    cs.CVarXiv:2608.30521v12026
  5. A Sanity Check for AI-generated Image Detection

    Shilin Yan, Ouxiang Li, Jiayin Cai +4

    cs.CVarXiv:2406.19435v32024
  6. 3D Visual Perception for Self-Driving Cars using a Multi-Camera System: Calibration, Mapping, Localization, and Obstacle Detection

    Christian Häne, Lionel Heng, Gim Hee Lee +4

    cs.CVarXiv:1708.09839v12017
  7. FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection

    Shaoqing Xu, Dingfu Zhou, Jin Fang +3

    cs.CVarXiv:2106.12449v22021
  8. Diffusion Probabilistic Models beat GANs on Medical Images

    Gustav Müller-Franzes, Jan Moritz Niehues, Firas Khader +8

    eess.IVcs.CVarXiv:2212.07501v12022
  9. SuperDepth: Self-Supervised, Super-Resolved Monocular Depth Estimation

    Sudeep Pillai, Rares Ambrus, Adrien Gaidon

    cs.CVcs.AIcs.LGarXiv:1810.01849v12018
  10. Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification

    Renchun You, Zhiyao Guo, Lei Cui +3

    cs.CVarXiv:1912.07872v22019
  11. Normalized and Geometry-Aware Self-Attention Network for Image Captioning

    Longteng Guo, Jing Liu, Xinxin Zhu +3

    cs.CVcs.CLcs.MMarXiv:2003.08897v12020
  12. A Style-Aware Content Loss for Real-time HD Style Transfer

    Artsiom Sanakoyeu, Dmytro Kotovenko, Sabine Lang +1

    cs.CVarXiv:1807.10201v22018
  13. Details or Artifacts: A Locally Discriminative Learning Approach to Realistic Image Super-Resolution

    Jie Liang, Hui Zeng, Lei Zhang

    eess.IVcs.CVarXiv:2203.09195v12022
  14. GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

    Yang Chen, Canyu Shen, Xinzhe Rao +5

    cs.CVarXiv:2608.29113v12026
  15. Apple Flower Detection using Deep Convolutional Networks

    Philipe A. Dias, Amy Tabb, Henry Medeiros

    cs.CVarXiv:1809.06357v12018
  16. Towards Crafting Text Adversarial Samples

    Suranjana Samanta, Sameep Mehta

    cs.LGcs.AIcs.CLarXiv:1707.02812v12017
  17. Deep Edge Guided Recurrent Residual Learning for Image Super-Resolution

    Wenhan Yang, Jiashi Feng, Jianchao Yang +4

    cs.CVarXiv:1604.08671v22016
  18. Attention Correctness in Neural Image Captioning

    Chenxi Liu, Junhua Mao, Fei Sha +1

    cs.CVcs.CLcs.LGarXiv:1605.09553v22016
  19. Person Re-Identification by Discriminative Selection in Video Ranking

    Taiqing Wang, Shaogang Gong, Xiatian Zhu +1

    cs.CVarXiv:1601.06260v12016
  20. Image Restoration using Total Variation Regularized Deep Image Prior

    Jiaming Liu, Yu Sun, Xiaojian Xu +1

    cs.CVarXiv:1810.12864v12018
  21. A Benchmark for Studying Diabetic Retinopathy: Segmentation, Grading, and Transferability

    Yi Zhou, Boyang Wang, Lei Huang +2

    cs.CVarXiv:2008.09772v32020
  22. Training Quantized Nets: A Deeper Understanding

    Hao Li, Soham De, Zheng Xu +3

    cs.LGcs.CVstat.MLarXiv:1706.02379v32017
  23. Unsupervised Generative Adversarial Cross-modal Hashing

    Jian Zhang, Yuxin Peng, Mingkuan Yuan

    cs.CVarXiv:1712.00358v12017
  24. One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic Control

    Wenlong Huang, Igor Mordatch, Deepak Pathak

    cs.LGcs.CVstat.MLarXiv:2007.04976v12020
  25. Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning

    Jian Ma, Junhao Liang, Chen Chen +1

    cs.CVarXiv:2307.11410v22023
  26. Knowledge Matters: Radiology Report Generation with General and Specific Knowledge

    Shuxin Yang, Xian Wu, Shen Ge +2

    eess.IVcs.CLcs.CVarXiv:2112.15009v22021
  27. An End-to-End Compression Framework Based on Convolutional Neural Networks

    Feng Jiang, Wen Tao, Shaohui Liu +3

    cs.CVarXiv:1708.00838v12017
  28. LightFuse: Relightable Interactive Gaussian Scene Reconstruction via Multi-Scan Fusion and 2D Gaussian Ray Tracing

    Haonan Zhou, Gaoxiang Linghu, Youlin Jia +5

    cs.CVarXiv:2608.29269v12026
  29. FISICA: A Deployed Service for Plantar-Pressure and Posture Assessment with Ontology-Grounded Recommendation

    Juhwan Song, Heejung Kim, Juntae Noh +5

    cs.CVcs.HCcs.IRarXiv:2608.29336v12026
  30. MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection

    Haoyang He, Yuhu Bai, Jiangning Zhang +7

    cs.CVarXiv:2404.06564v42024
  31. Color Constancy Using CNNs

    Simone Bianco, Claudio Cusano, Raimondo Schettini

    cs.CVarXiv:1504.04548v12015
  32. Diffusion-based Image Translation using Disentangled Style and Content Representation

    Gihyun Kwon, Jong Chul Ye

    cs.CVcs.AIcs.LGarXiv:2209.15264v22022
  33. PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

    Qitao Zhao, Ce Zheng, Mengyuan Liu +2

    cs.CVarXiv:2303.17472v12023
  34. Min-Entropy Latent Model for Weakly Supervised Object Detection

    Fang Wan, Pengxu Wei, Zhenjun Han +2

    cs.CVarXiv:1902.06057v12019
  35. Semantic Image Synthesis via Diffusion Models

    Wengang Zhou, Weilun Wang, Jianmin Bao +4

    cs.CVarXiv:2207.00050v42022
  36. Explainable Artificial Intelligence (XAI) in Computational Pathology: Definitions, Taxonomy, and Recommendations

    Shubham Innani, Suhang You, Adam Shephard +13

    cs.AIcs.CVarXiv:2608.28820v12026
  37. Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation

    Tianrui Hui, Shaofei Huang, Qisong Han +6

    cs.CVarXiv:2608.29121v12026
  38. Cross-View Image Synthesis using Conditional GANs

    Krishna Regmi, Ali Borji

    cs.CVarXiv:1803.03396v22018
  39. Rigid and Articulated Point Registration with Expectation Conditional Maximization

    Radu Horaud, Florence Forbes, Manuel Yguel +2

    cs.CVarXiv:2012.05191v12020
  40. DCFNet: Discriminant Correlation Filters Network for Visual Tracking

    Qiang Wang, Jin Gao, Junliang Xing +2

    cs.CVarXiv:1704.04057v12017
  41. SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes

    Yutao Cui, Chenkai Zeng, Xiaoyu Zhao +3

    cs.CVarXiv:2304.05170v22023
  42. Automatically Discovering and Learning New Visual Categories with Ranking Statistics

    Kai Han, Sylvestre-Alvise Rebuffi, Sebastien Ehrhardt +2

    cs.CVarXiv:2002.05714v12020
  43. Task-Driven Modular Networks for Zero-Shot Compositional Learning

    Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta +1

    cs.CVarXiv:1905.05908v12019
  44. ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views

    Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski +1

    cs.CVarXiv:2608.28895v12026
  45. Distributed Deep Learning Model for Intelligent Video Surveillance Systems with Edge Computing

    Jianguo Chen, Kenli Li, Qingying Deng +2

    cs.CVcs.LGarXiv:1904.06400v12019
  46. Frequency-Assisted Mamba for Remote Sensing Image Super-Resolution

    Yi Xiao, Qiangqiang Yuan, Kui Jiang +3

    cs.CVarXiv:2405.04964v22024
  47. Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

    Yiting Qu, Xinyue Shen, Xinlei He +3

    cs.CVcs.CRcs.CYarXiv:2305.13873v22023
  48. Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut

    Yangtao Wang, Xi Shen, Shell Hu +3

    cs.CVstat.MLarXiv:2202.11539v22022
  49. Fully Convolutional Grasp Detection Network with Oriented Anchor Box

    Xinwen Zhou, Xuguang Lan, Hanbo Zhang +3

    cs.ROcs.CVarXiv:1803.02209v12018
  50. Learnable Manifold Alignment (LeMA) : A Semi-supervised Cross-modality Learning Framework for Land Cover and Land Use Classification

    Danfeng Hong, Naoto Yokoya, Nan Ge +2

    cs.CVarXiv:1901.02838v12019
  51. Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing

    Yapeng Tian, Dingzeyu Li, Chenliang Xu

    cs.CVcs.MMcs.SDarXiv:2007.10558v12020
  52. Machine Learning Information Fusion in Earth Observation: A Comprehensive Review of Methods, Applications and Data Sources

    S. Salcedo-Sanz, P. Ghamisi, M. Piles +7

    cs.CVcs.LGarXiv:2012.05795v12020
  53. Vision-Language Models in Remote Sensing: Current Progress and Future Trends

    Xiang Li, Congcong Wen, Yuan Hu +2

    cs.CVcs.AIarXiv:2305.05726v22023
  54. A Perspective Analysis of Handwritten Signature Technology

    Moises Diaz, Miguel A. Ferrer, Donato Impedovo +3

    cs.CVarXiv:2405.13555v12024
  55. MWIR-4-Plastic: The Identification of Complex End-of-Life Industrial Plastic using Mid-wave Infrared Hyperspectral Imaging and Machine Learning

    Elias Arbash, Andréa de Lima Ribeiro, Filipa Simões +9

    cs.CVcs.LGarXiv:2608.28874v12026
  56. Domain Generalization for Medical Imaging Classification with Linear-Dependency Regularization

    Haoliang Li, YuFei Wang, Renjie Wan +3

    cs.CVcs.LGeess.IVarXiv:2009.12829v32020
  57. Video Probabilistic Diffusion Models in Projected Latent Space

    Sihyun Yu, Kihyuk Sohn, Subin Kim +1

    cs.CVcs.LGarXiv:2302.07685v22023
  58. SinNeRF: Training Neural Radiance Fields on Complex Scenes from a Single Image

    Dejia Xu, Yifan Jiang, Peihao Wang +3

    cs.CVarXiv:2204.00928v22022
  59. Deep Optics for Monocular Depth Estimation and 3D Object Detection

    Julie Chang, Gordon Wetzstein

    cs.CVeess.IVarXiv:1904.08601v12019
  60. Omni-frequency Channel-selection Representations for Unsupervised Anomaly Detection

    Yufei Liang, Jiangning Zhang, Shiwei Zhao +3

    cs.CVarXiv:2203.00259v22022