Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,021 to 13,080 of 18,866

  1. RUBi: Reducing Unimodal Biases in Visual Question Answering

    Remi Cadene, Corentin Dancette, Hedi Ben-younes +2

    cs.CVcs.CLcs.LGarXiv:1906.10169v22019
  2. CrossFuse: A Novel Cross Attention Mechanism based Infrared and Visible Image Fusion Approach

    Hui Li, Xiao-Jun Wu

    cs.CVarXiv:2406.10581v12024
  3. FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding

    Dian Shao, Yue Zhao, Bo Dai +1

    cs.CVarXiv:2004.06704v12020
  4. Learning Monocular Reactive UAV Control in Cluttered Natural Environments

    Stephane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar +4

    cs.ROcs.CVcs.LGarXiv:1211.1690v12012
  5. Multi-view Low-rank Sparse Subspace Clustering

    Maria Brbic, Ivica Kopriva

    cs.CVcs.LGmath.OCarXiv:1708.08732v12017
  6. Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural Networks

    Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch +1

    cs.CVcs.LGarXiv:1810.08468v12018
  7. Boosting Few-Shot Visual Learning with Self-Supervision

    Spyros Gidaris, Andrei Bursuc, Nikos Komodakis +2

    cs.CVcs.LGarXiv:1906.05186v12019
  8. Learning Steerable Filters for Rotation Equivariant CNNs

    Maurice Weiler, Fred A. Hamprecht, Martin Storath

    cs.LGcs.CVarXiv:1711.07289v32017
  9. Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise

    Arpit Bansal, Eitan Borgnia, Hong-Min Chu +6

    cs.CVcs.LGarXiv:2208.09392v12022
    Summaries:한국어
  10. Rethinking and Improving Relative Position Encoding for Vision Transformer

    Kan Wu, Houwen Peng, Minghao Chen +2

    cs.CVarXiv:2107.14222v12021
  11. AutoZOOM: Autoencoder-based Zeroth Order Optimization Method for Attacking Black-box Neural Networks

    Chun-Chen Tu, Paishun Ting, Pin-Yu Chen +5

    cs.CVcs.CRstat.MLarXiv:1805.11770v52018
  12. A Fast and Accurate One-Stage Approach to Visual Grounding

    Zhengyuan Yang, Boqing Gong, Liwei Wang +3

    cs.CVarXiv:1908.06354v12019
  13. No-Reference Image Quality Assessment via Transformers, Relative Ranking, and Self-Consistency

    S. Alireza Golestaneh, Saba Dadsetan, Kris M. Kitani

    eess.IVcs.CVarXiv:2108.06858v22021
  14. SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis

    Mengqi Ji, Juergen Gall, Haitian Zheng +2

    cs.CVarXiv:1708.01749v12017
  15. Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff

    Yochai Blau, Tomer Michaeli

    cs.LGcs.CVcs.ITarXiv:1901.07821v42019
  16. Importance Weighted Adversarial Nets for Partial Domain Adaptation

    Jing Zhang, Zewei Ding, Wanqing Li +1

    cs.CVarXiv:1803.09210v22018
  17. Fast Face-swap Using Convolutional Neural Networks

    Iryna Korshunova, Wenzhe Shi, Joni Dambre +1

    cs.CVarXiv:1611.09577v22016
  18. From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality

    Zhenqiang Ying, Haoran Niu, Praful Gupta +3

    cs.CVcs.MMeess.IVarXiv:1912.10088v12019
  19. Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

    Ji Soo Lee, Jinyoung Park, Seohyun Lee +4

    cs.CVarXiv:2608.26684v12026
  20. HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

    Fei Ma, Zebang Cheng, Minghui Li +11

    cs.CVarXiv:2608.26517v12026
  21. A General Optimization-based Framework for Global Pose Estimation with Multiple Sensors

    Tong Qin, Shaozu Cao, Jie Pan +1

    cs.CVarXiv:1901.03642v12019
  22. GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image

    Mingjian Zhu, Hanting Chen, Qiangyu Yan +7

    cs.CVarXiv:2306.08571v22023
  23. Learning Woody Clearing With Loss Alignment for Zero-Shot Regrowth and Woody Segmentation

    Kal Backman, Jared Wood, Adam Roff

    cs.CVarXiv:2608.26489v12026
  24. Hyperspectral Image Denoising Employing a Spatial-Spectral Deep Residual Convolutional Neural Network

    Qiangqiang Yuan, Qiang Zhang, Jie Li +2

    cs.CVarXiv:1806.00183v32018
  25. Text2Mesh: Text-Driven Neural Stylization for Meshes

    Oscar Michel, Roi Bar-On, Richard Liu +2

    cs.CVcs.CLcs.GRarXiv:2112.03221v12021
  26. From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation

    Haowen Gu, Gensheng Pei, Junzhu Mao +3

    cs.CVcs.AIarXiv:2608.26856v12026
    Summaries:简体中文
  27. DiffIR: Efficient Diffusion Model for Image Restoration

    Bin Xia, Yulun Zhang, Shiyin Wang +5

    cs.CVarXiv:2303.09472v32023
  28. A survey of the Vision Transformers and their CNN-Transformer based Variants

    Asifullah Khan, Zunaira Rauf, Anabia Sohail +4

    cs.CVarXiv:2305.09880v42023
  29. RECAP-Forcing: Retaining Content Appearances for Long Video Generation

    Haiyang Xu, Zheng Ding, Zhuowen Tu

    cs.CVarXiv:2608.26671v12026
  30. MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures

    Zhiqin Chen, Thomas Funkhouser, Peter Hedman +1

    cs.CVcs.GRcs.LGarXiv:2208.00277v52022
  31. Disentangled Person Image Generation

    Liqian Ma, Qianru Sun, Stamatios Georgoulis +3

    cs.CVarXiv:1712.02621v42017
  32. FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces

    Andreas Rössler, Davide Cozzolino, Luisa Verdoliva +3

    cs.CVarXiv:1803.09179v12018
  33. MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis

    Jiancheng Yang, Rui Shi, Bingbing Ni

    cs.CVcs.AIcs.LGarXiv:2010.14925v42020
  34. PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting

    Kai Zhang, Fujun Luan, Qianqian Wang +2

    cs.CVcs.GRarXiv:2104.00674v12021
  35. Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

    Shihao Wang, Yingfei Liu, Tiancai Wang +2

    cs.CVarXiv:2303.11926v22023
  36. GRNet: Gridding Residual Network for Dense Point Cloud Completion

    Haozhe Xie, Hongxun Yao, Shangchen Zhou +3

    cs.CVcs.LGeess.IVarXiv:2006.03761v42020
  37. COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

    Yansong Tang, Dajun Ding, Yongming Rao +5

    cs.CVarXiv:1903.02874v12019
  38. CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent Reconstruction

    Jean-Daniel de Ambrogi, Aladine Chetouani, Vincent Nguyen +1

    cs.CVcs.ROarXiv:2608.26868v12026
  39. Plug-and-Play Methods Provably Converge with Properly Trained Denoisers

    Ernest K. Ryu, Jialin Liu, Sicheng Wang +3

    cs.CVeess.IVarXiv:1905.05406v12019
  40. G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

    Zehua Hao, Fang Liu, Qinliang Wang +3

    cs.CVarXiv:2608.26744v12026
  41. Analyzing and Improving the Training Dynamics of Diffusion Models

    Tero Karras, Miika Aittala, Jaakko Lehtinen +3

    cs.CVcs.AIcs.LGarXiv:2312.02696v22023
  42. Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping

    Michael Xie, Neal Jean, Marshall Burke +2

    cs.CVcs.CYarXiv:1510.00098v22015
  43. STDP-based spiking deep convolutional neural networks for object recognition

    Saeed Reza Kheradpisheh, Mohammad Ganjtabesh, Simon J Thorpe +1

    cs.CVarXiv:1611.01421v32016
  44. PhysDiff: Physics-Guided Human Motion Diffusion Model

    Ye Yuan, Jiaming Song, Umar Iqbal +2

    cs.CVcs.AIcs.GRarXiv:2212.02500v32022
  45. SMIL: Multimodal Learning with Severely Missing Modality

    Mengmeng Ma, Jian Ren, Long Zhao +3

    cs.CVarXiv:2103.05677v12021
  46. PandaGPT: One Model To Instruction-Follow Them All

    Yixuan Su, Tian Lan, Huayang Li +3

    cs.CLcs.CVarXiv:2305.16355v12023
  47. Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting

    Shuo Li, Mike Davies, Mehrdad Yaghoobi

    cs.CVcs.LGarXiv:2608.26812v12026
  48. Deep Convolutional Ranking for Multilabel Image Annotation

    Yunchao Gong, Yangqing Jia, Thomas Leung +2

    cs.CVarXiv:1312.4894v22013
  49. PyHST2: an hybrid distributed code for high speed tomographic reconstruction with iterative reconstruction and a priori knowledge capabilities

    Alessandro Mirone, Emmanuelle Gouillart, Emmanuel Brun +2

    math.NAcs.CVarXiv:1306.1392v12013
  50. Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation

    Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang +4

    cs.CVarXiv:2007.09183v12020
  51. Population Structure Analysis of an Inbred Population using Quantitative Shape Phenotyping from Stereo Retinal Photographs

    Li Tang, Michael D Abramoff

    cs.LGcs.CVq-bio.QMarXiv:2608.15471v12026
  52. Video-FLAIR: Not Whether to Reason, But How

    Yogesh Kulkarni, Pooyan Fazli

    cs.CVarXiv:2608.26495v12026
  53. Synthetic Data from Diffusion Models Improves ImageNet Classification

    Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia +2

    cs.CVcs.AIcs.CLarXiv:2304.08466v12023
  54. Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator

    Kumju Jo, Heesun Jung, Sungyong Baik

    cs.CVarXiv:2608.26624v12026
  55. Gaussian YOLOv3: An Accurate and Fast Object Detector Using Localization Uncertainty for Autonomous Driving

    Jiwoong Choi, Dayoung Chun, Hyun Kim +1

    cs.CVarXiv:1904.04620v22019
    Summaries:한국어
  56. Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction

    Chen-Hsuan Lin, Chen Kong, Simon Lucey

    cs.CVcs.LGarXiv:1706.07036v12017
  57. PP-YOLOE: An evolved version of YOLO

    Shangliang Xu, Xinxin Wang, Wenyu Lv +8

    cs.CVarXiv:2203.16250v32022
  58. Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

    Jiaming Zhou, Qihang Zhang, Gangwei Xu +9

    cs.ROcs.CVarXiv:2608.26103v22026
    Summaries:한국어
  59. A Discriminative CNN Video Representation for Event Detection

    Zhongwen Xu, Yi Yang, Alexander G. Hauptmann

    cs.CVarXiv:1411.4006v12014
    Summaries:한국어
  60. Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

    Pan Lu, Liang Qiu, Kai-Wei Chang +5

    cs.LGcs.AIcs.CLarXiv:2209.14610v32022