Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,241 to 15,300 of 18,866
Person Re-identification in the Wild
Liang Zheng, Hengheng Zhang, Shaoyan Sun +3
cs.CVarXiv:1604.02531v22016AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
Ruilong Li, Shan Yang, David A. Ross +1
cs.CVcs.GRcs.MMarXiv:2101.08779v32021ORBIT++: Benchmarking SfM in the Wild with 360° Video
Sara Sabour, Linyi Jin, Richard Tucker +9
cs.CVarXiv:2608.22039v12026Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting
Shuangkang Fang, I-Chao Shen, Xuanyang Zhang +5
cs.CVarXiv:2602.20933v12026Towards Fast Computation of Certified Robustness for ReLU Networks
Tsui-Wei Weng, Huan Zhang, Hongge Chen +5
stat.MLcs.CRcs.CVarXiv:1804.09699v42018$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
Siting Wang, Xiaofeng Wang, Zheng Zhu +7
cs.ROcs.CVarXiv:2603.02083v22026Learning to Propagate Labels: Transductive Propagation Network for Few-shot Learning
Yanbin Liu, Juho Lee, Minseop Park +4
cs.LGcs.CVcs.NEarXiv:1805.10002v52018FlyPose: Towards Robust Human Pose Estimation From Aerial Views
Hassaan Farooq, Marvin Brenner, Peter Stütz
cs.CVcs.ROarXiv:2601.05747v22026VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR
Yani Guan, Dengpan Dong, Shuang Luo +6
cs.CVcs.IRcs.LGarXiv:2608.22183v12026On-Policy Self-Distillation in Diffusion Models
Wei Zhou, Xiongwei Zhu, Lingdong Kong +14
cs.CVarXiv:2608.24646v120263D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
Hao Tang, Ting Huang, Zeyu Zhang
cs.CVarXiv:2601.06496v12026UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao +9
cs.CVarXiv:2601.11522v12026GutenOCR: A Grounded Vision-Language Front-End for Documents
Hunter Heidenreich, Ben Elliott, Olivia Dinica +1
cs.CVcs.AIcs.CLarXiv:2601.14490v22026Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
Andrew Slavin Ross, Finale Doshi-Velez
cs.LGcs.CRcs.CVarXiv:1711.09404v12017NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
Zhuchenyang Liu, Yao Zhang, Yu Xiao
cs.IRcs.CVcs.LGarXiv:2603.12824v22026ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
cs.CVarXiv:1906.02467v12019Masked Autoencoders for Point Cloud Self-supervised Learning
Yatian Pang, Wenxiao Wang, Francis E. H. Tay +3
cs.CVarXiv:2203.06604v22022Urban Socio-Semantic Segmentation with Vision-Language Reasoning
Yu Wang, Yi Wang, Rui Dai +4
cs.CVcs.AIcs.CYarXiv:2601.10477v22026Enhancing Underwater Imagery using Generative Adversarial Networks
Cameron Fabbri, Md Jahidul Islam, Junaed Sattar
cs.CVcs.ROarXiv:1801.04011v12018nnFormer: Interleaved Transformer for Volumetric Segmentation
Hong-Yu Zhou, Jiansen Guo, Yinghao Zhang +3
cs.CVarXiv:2109.03201v62021VideoMaMa: Mask-Guided Video Matting via Generative Prior
Sangbeom Lim, Seoung Wug Oh, Jiahui Huang +3
cs.CVcs.AIarXiv:2601.14255v12026Gated Context Aggregation Network for Image Dehazing and Deraining
Dongdong Chen, Mingming He, Qingnan Fan +5
cs.CVarXiv:1811.08747v22018Structure and Content-Guided Video Synthesis with Diffusion Models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian +2
cs.CVarXiv:2302.03011v12023A Comprehensive Overhaul of Feature Distillation
Byeongho Heo, Jeesoo Kim, Sangdoo Yun +3
cs.CVcs.LGarXiv:1904.01866v22019Learning a Deep Embedding Model for Zero-Shot Learning
Li Zhang, Tao Xiang, Shaogang Gong
cs.CVarXiv:1611.05088v42016GroupViT: Semantic Segmentation Emerges from Text Supervision
Jiarui Xu, Shalini De Mello, Sifei Liu +4
cs.CVarXiv:2202.11094v52022VIOLA: Towards Video In-Context Learning with Minimal Annotations
Ryo Fujii, Hideo Saito, Ryo Hachiuma
cs.CVcs.AIarXiv:2601.15549v12026GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
Shuguang Zhang, Junhong Lian, Guoxin Yu +2
cs.CVcs.AIcs.CLarXiv:2601.20618v12026Deep Learning COVID-19 Features on CXR using Limited Training Data Sets
Yujin Oh, Sangjoon Park, Jong Chul Ye
eess.IVcs.CVcs.LGarXiv:2004.05758v22020Deep Pyramidal Residual Networks
Dongyoon Han, Jiwhan Kim, Junmo Kim
cs.CVarXiv:1610.02915v42016Visual Transformers: Token-based Image Representation and Processing for Computer Vision
Bichen Wu, Chenfeng Xu, Xiaoliang Dai +7
cs.CVcs.LGeess.IVarXiv:2006.03677v42020RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal +7
cs.LGcs.AIcs.CLarXiv:2304.06767v42023PLANING: A Loosely Coupled Triangle-Gaussian Framework for Streaming 3D Reconstruction
Changjian Jiang, Kerui Ren, Xudong Li +8
cs.CVarXiv:2601.22046v42026Medical Image Synthesis with Context-Aware Generative Adversarial Networks
Dong Nie, Roger Trullo, Caroline Petitjean +2
cs.CVarXiv:1612.05362v12016Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
Aditya Prakash, Kashyap Chitta, Andreas Geiger
cs.CVcs.AIcs.LGarXiv:2104.09224v12021Variational Information Distillation for Knowledge Transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou +2
cs.CVcs.AIcs.LGarXiv:1904.05835v12019Glance and Focus Reinforcement for Pan-cancer Screening
Linshan Wu, Jiaxin Zhuang, Hao Chen
cs.CVarXiv:2601.19103v22026Big Self-Supervised Models Advance Medical Image Classification
Shekoofeh Azizi, Basil Mustafa, Fiona Ryan +9
eess.IVcs.CVcs.LGarXiv:2101.05224v22021VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation
Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah
cs.CVarXiv:2608.22521v12026Revisiting Self-Supervised Visual Representation Learning
Alexander Kolesnikov, Xiaohua Zhai, Lucas Beyer
cs.CVarXiv:1901.09005v12019Conformer: Local Features Coupling Global Representations for Visual Recognition
Zhiliang Peng, Wei Huang, Shanzhi Gu +4
cs.CVarXiv:2105.03889v12021VISTA-PATH: An interactive foundation model for pathology image segmentation and quantitative analysis in computational pathology
Peixian Liang, Songhao Li, Shunsuke Koga +5
cs.CVarXiv:2601.16451v12026Rethinking Spatial Dimensions of Vision Transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han +3
cs.CVarXiv:2103.16302v22021WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
Hui Zhang, Juntao Liu, Zongkai Liu +4
cs.CVarXiv:2603.11593v22026Emu3: Next-Token Prediction is All You Need
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22
cs.CVarXiv:2409.18869v12024SimpleGPT: Improving GPT via A Simple Normalization Strategy
Marco Chen, Xianbiao Qi, Yelin He +2
cs.LGcs.CLcs.CVarXiv:2602.01212v12026A Survey of Appearance Models in Visual Object Tracking
Xi Li, Weiming Hu, Chunhua Shen +3
cs.CVarXiv:1303.4803v12013Dynamic Memory Networks for Visual and Textual Question Answering
Caiming Xiong, Stephen Merity, Richard Socher
cs.NEcs.CLcs.CVarXiv:1603.01417v12016Toward Real-World Single Image Super-Resolution: A New Benchmark and A New Model
Jianrui Cai, Hui Zeng, Hongwei Yong +2
cs.CVarXiv:1904.00523v12019Bridge Damage Detection from Low-Light UAV Imagery via Degradation-Aware Mixture-of-Experts Enhancement
Hu Wang, Hongxu Pu, Zhiqi Hu +2
cs.CVarXiv:2608.23136v12026VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
Hassan Akbari, Liangzhe Yuan, Rui Qian +4
cs.CVcs.AIcs.LGarXiv:2104.11178v32021What makes for effective detection proposals?
Jan Hosang, Rodrigo Benenson, Piotr Dollár +1
cs.CVarXiv:1502.05082v32015Multi-scale Interactive Network for Salient Object Detection
Youwei Pang, Xiaoqi Zhao, Lihe Zhang +1
cs.CVarXiv:2007.09062v12020WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching
Yihan Wang, Jia Deng
cs.CVarXiv:2603.24836v32026Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
Youliang Zhang, Zhengguang Zhou, Zhentao Yu +11
cs.CVcs.AIcs.CLarXiv:2602.01538v12026Enhancing Multi-Image Understanding through Delimiter Token Scaling
Minyoung Lee, Yeji Park, Dongjun Hwang +3
cs.CVarXiv:2602.01984v22026SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation
Yang Zou, Jongheon Jeong, Latha Pemula +2
cs.CVarXiv:2207.14315v12022Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar
Taqi Hamoda, Nuno Gracias
cs.CVarXiv:2608.23479v12026BlendedMVS: A Large-scale Dataset for Generalized Multi-view Stereo Networks
Yao Yao, Zixin Luo, Shiwei Li +5
cs.CVarXiv:1911.10127v22019What Remains Normal? Clean Images Miss Useful Near-Defect Normal Patches for Anomaly Detection
Joongwon Chae, Runming Wang, Peiwu Qin
cs.CVarXiv:2608.23299v12026