Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,021 to 13,080 of 18,866
RUBi: Reducing Unimodal Biases in Visual Question Answering
Remi Cadene, Corentin Dancette, Hedi Ben-younes +2
cs.CVcs.CLcs.LGarXiv:1906.10169v22019CrossFuse: A Novel Cross Attention Mechanism based Infrared and Visible Image Fusion Approach
Hui Li, Xiao-Jun Wu
cs.CVarXiv:2406.10581v12024FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding
Dian Shao, Yue Zhao, Bo Dai +1
cs.CVarXiv:2004.06704v12020Learning Monocular Reactive UAV Control in Cluttered Natural Environments
Stephane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar +4
cs.ROcs.CVcs.LGarXiv:1211.1690v12012Multi-view Low-rank Sparse Subspace Clustering
Maria Brbic, Ivica Kopriva
cs.CVcs.LGmath.OCarXiv:1708.08732v12017Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural Networks
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch +1
cs.CVcs.LGarXiv:1810.08468v12018Boosting Few-Shot Visual Learning with Self-Supervision
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis +2
cs.CVcs.LGarXiv:1906.05186v12019Learning Steerable Filters for Rotation Equivariant CNNs
Maurice Weiler, Fred A. Hamprecht, Martin Storath
cs.LGcs.CVarXiv:1711.07289v32017Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu +6
cs.CVcs.LGarXiv:2208.09392v12022Summaries:한국어Rethinking and Improving Relative Position Encoding for Vision Transformer
Kan Wu, Houwen Peng, Minghao Chen +2
cs.CVarXiv:2107.14222v12021AutoZOOM: Autoencoder-based Zeroth Order Optimization Method for Attacking Black-box Neural Networks
Chun-Chen Tu, Paishun Ting, Pin-Yu Chen +5
cs.CVcs.CRstat.MLarXiv:1805.11770v52018A Fast and Accurate One-Stage Approach to Visual Grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang +3
cs.CVarXiv:1908.06354v12019No-Reference Image Quality Assessment via Transformers, Relative Ranking, and Self-Consistency
S. Alireza Golestaneh, Saba Dadsetan, Kris M. Kitani
eess.IVcs.CVarXiv:2108.06858v22021SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis
Mengqi Ji, Juergen Gall, Haitian Zheng +2
cs.CVarXiv:1708.01749v12017Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff
Yochai Blau, Tomer Michaeli
cs.LGcs.CVcs.ITarXiv:1901.07821v42019Importance Weighted Adversarial Nets for Partial Domain Adaptation
Jing Zhang, Zewei Ding, Wanqing Li +1
cs.CVarXiv:1803.09210v22018Fast Face-swap Using Convolutional Neural Networks
Iryna Korshunova, Wenzhe Shi, Joni Dambre +1
cs.CVarXiv:1611.09577v22016From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality
Zhenqiang Ying, Haoran Niu, Praful Gupta +3
cs.CVcs.MMeess.IVarXiv:1912.10088v12019Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs
Ji Soo Lee, Jinyoung Park, Seohyun Lee +4
cs.CVarXiv:2608.26684v12026HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence
Fei Ma, Zebang Cheng, Minghui Li +11
cs.CVarXiv:2608.26517v12026A General Optimization-based Framework for Global Pose Estimation with Multiple Sensors
Tong Qin, Shaozu Cao, Jie Pan +1
cs.CVarXiv:1901.03642v12019GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image
Mingjian Zhu, Hanting Chen, Qiangyu Yan +7
cs.CVarXiv:2306.08571v22023Learning Woody Clearing With Loss Alignment for Zero-Shot Regrowth and Woody Segmentation
Kal Backman, Jared Wood, Adam Roff
cs.CVarXiv:2608.26489v12026Hyperspectral Image Denoising Employing a Spatial-Spectral Deep Residual Convolutional Neural Network
Qiangqiang Yuan, Qiang Zhang, Jie Li +2
cs.CVarXiv:1806.00183v32018Text2Mesh: Text-Driven Neural Stylization for Meshes
Oscar Michel, Roi Bar-On, Richard Liu +2
cs.CVcs.CLcs.GRarXiv:2112.03221v12021From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Haowen Gu, Gensheng Pei, Junzhu Mao +3
cs.CVcs.AIarXiv:2608.26856v12026Summaries:简体中文DiffIR: Efficient Diffusion Model for Image Restoration
Bin Xia, Yulun Zhang, Shiyin Wang +5
cs.CVarXiv:2303.09472v32023A survey of the Vision Transformers and their CNN-Transformer based Variants
Asifullah Khan, Zunaira Rauf, Anabia Sohail +4
cs.CVarXiv:2305.09880v42023RECAP-Forcing: Retaining Content Appearances for Long Video Generation
Haiyang Xu, Zheng Ding, Zhuowen Tu
cs.CVarXiv:2608.26671v12026MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures
Zhiqin Chen, Thomas Funkhouser, Peter Hedman +1
cs.CVcs.GRcs.LGarXiv:2208.00277v52022Disentangled Person Image Generation
Liqian Ma, Qianru Sun, Stamatios Georgoulis +3
cs.CVarXiv:1712.02621v42017FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva +3
cs.CVarXiv:1803.09179v12018MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis
Jiancheng Yang, Rui Shi, Bingbing Ni
cs.CVcs.AIcs.LGarXiv:2010.14925v42020PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting
Kai Zhang, Fujun Luan, Qianqian Wang +2
cs.CVcs.GRarXiv:2104.00674v12021Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
Shihao Wang, Yingfei Liu, Tiancai Wang +2
cs.CVarXiv:2303.11926v22023GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou +3
cs.CVcs.LGeess.IVarXiv:2006.03761v42020COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
Yansong Tang, Dajun Ding, Yongming Rao +5
cs.CVarXiv:1903.02874v12019CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent Reconstruction
Jean-Daniel de Ambrogi, Aladine Chetouani, Vincent Nguyen +1
cs.CVcs.ROarXiv:2608.26868v12026Plug-and-Play Methods Provably Converge with Properly Trained Denoisers
Ernest K. Ryu, Jialin Liu, Sicheng Wang +3
cs.CVeess.IVarXiv:1905.05406v12019G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification
Zehua Hao, Fang Liu, Qinliang Wang +3
cs.CVarXiv:2608.26744v12026Analyzing and Improving the Training Dynamics of Diffusion Models
Tero Karras, Miika Aittala, Jaakko Lehtinen +3
cs.CVcs.AIcs.LGarXiv:2312.02696v22023Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping
Michael Xie, Neal Jean, Marshall Burke +2
cs.CVcs.CYarXiv:1510.00098v22015STDP-based spiking deep convolutional neural networks for object recognition
Saeed Reza Kheradpisheh, Mohammad Ganjtabesh, Simon J Thorpe +1
cs.CVarXiv:1611.01421v32016PhysDiff: Physics-Guided Human Motion Diffusion Model
Ye Yuan, Jiaming Song, Umar Iqbal +2
cs.CVcs.AIcs.GRarXiv:2212.02500v32022SMIL: Multimodal Learning with Severely Missing Modality
Mengmeng Ma, Jian Ren, Long Zhao +3
cs.CVarXiv:2103.05677v12021PandaGPT: One Model To Instruction-Follow Them All
Yixuan Su, Tian Lan, Huayang Li +3
cs.CLcs.CVarXiv:2305.16355v12023Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting
Shuo Li, Mike Davies, Mehrdad Yaghoobi
cs.CVcs.LGarXiv:2608.26812v12026Deep Convolutional Ranking for Multilabel Image Annotation
Yunchao Gong, Yangqing Jia, Thomas Leung +2
cs.CVarXiv:1312.4894v22013PyHST2: an hybrid distributed code for high speed tomographic reconstruction with iterative reconstruction and a priori knowledge capabilities
Alessandro Mirone, Emmanuelle Gouillart, Emmanuel Brun +2
math.NAcs.CVarXiv:1306.1392v12013Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang +4
cs.CVarXiv:2007.09183v12020Population Structure Analysis of an Inbred Population using Quantitative Shape Phenotyping from Stereo Retinal Photographs
Li Tang, Michael D Abramoff
cs.LGcs.CVq-bio.QMarXiv:2608.15471v12026Video-FLAIR: Not Whether to Reason, But How
Yogesh Kulkarni, Pooyan Fazli
cs.CVarXiv:2608.26495v12026Synthetic Data from Diffusion Models Improves ImageNet Classification
Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia +2
cs.CVcs.AIcs.CLarXiv:2304.08466v12023Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator
Kumju Jo, Heesun Jung, Sungyong Baik
cs.CVarXiv:2608.26624v12026Gaussian YOLOv3: An Accurate and Fast Object Detector Using Localization Uncertainty for Autonomous Driving
Jiwoong Choi, Dayoung Chun, Hyun Kim +1
cs.CVarXiv:1904.04620v22019Summaries:한국어Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction
Chen-Hsuan Lin, Chen Kong, Simon Lucey
cs.CVcs.LGarXiv:1706.07036v12017PP-YOLOE: An evolved version of YOLO
Shangliang Xu, Xinxin Wang, Wenyu Lv +8
cs.CVarXiv:2203.16250v32022Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Jiaming Zhou, Qihang Zhang, Gangwei Xu +9
cs.ROcs.CVarXiv:2608.26103v22026Summaries:한국어A Discriminative CNN Video Representation for Event Detection
Zhongwen Xu, Yi Yang, Alexander G. Hauptmann
cs.CVarXiv:1411.4006v12014Summaries:한국어Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Pan Lu, Liang Qiu, Kai-Wei Chang +5
cs.LGcs.AIcs.CLarXiv:2209.14610v32022