Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,241 to 15,300 of 18,866

  1. Person Re-identification in the Wild

    Liang Zheng, Hengheng Zhang, Shaoyan Sun +3

    cs.CVarXiv:1604.02531v22016
  2. AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

    Ruilong Li, Shan Yang, David A. Ross +1

    cs.CVcs.GRcs.MMarXiv:2101.08779v32021
  3. ORBIT++: Benchmarking SfM in the Wild with 360° Video

    Sara Sabour, Linyi Jin, Richard Tucker +9

    cs.CVarXiv:2608.22039v12026
  4. Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting

    Shuangkang Fang, I-Chao Shen, Xuanyang Zhang +5

    cs.CVarXiv:2602.20933v12026
  5. Towards Fast Computation of Certified Robustness for ReLU Networks

    Tsui-Wei Weng, Huan Zhang, Hongge Chen +5

    stat.MLcs.CRcs.CVarXiv:1804.09699v42018
  6. $π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

    Siting Wang, Xiaofeng Wang, Zheng Zhu +7

    cs.ROcs.CVarXiv:2603.02083v22026
  7. Learning to Propagate Labels: Transductive Propagation Network for Few-shot Learning

    Yanbin Liu, Juho Lee, Minseop Park +4

    cs.LGcs.CVcs.NEarXiv:1805.10002v52018
  8. FlyPose: Towards Robust Human Pose Estimation From Aerial Views

    Hassaan Farooq, Marvin Brenner, Peter Stütz

    cs.CVcs.ROarXiv:2601.05747v22026
  9. VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

    Yani Guan, Dengpan Dong, Shuang Luo +6

    cs.CVcs.IRcs.LGarXiv:2608.22183v12026
  10. On-Policy Self-Distillation in Diffusion Models

    Wei Zhou, Xiongwei Zhu, Lingdong Kong +14

    cs.CVarXiv:2608.24646v12026
  11. 3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence

    Hao Tang, Ting Huang, Zeyu Zhang

    cs.CVarXiv:2601.06496v12026
  12. UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation

    Ruiheng Zhang, Jingfeng Yao, Huangxuan Zhao +9

    cs.CVarXiv:2601.11522v12026
  13. GutenOCR: A Grounded Vision-Language Front-End for Documents

    Hunter Heidenreich, Ben Elliott, Olivia Dinica +1

    cs.CVcs.AIcs.CLarXiv:2601.14490v22026
  14. Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients

    Andrew Slavin Ross, Finale Doshi-Velez

    cs.LGcs.CRcs.CVarXiv:1711.09404v12017
  15. NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval

    Zhuchenyang Liu, Yao Zhang, Yu Xiao

    cs.IRcs.CVcs.LGarXiv:2603.12824v22026
  16. ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

    Zhou Yu, Dejing Xu, Jun Yu +4

    cs.CVarXiv:1906.02467v12019
  17. Masked Autoencoders for Point Cloud Self-supervised Learning

    Yatian Pang, Wenxiao Wang, Francis E. H. Tay +3

    cs.CVarXiv:2203.06604v22022
  18. Urban Socio-Semantic Segmentation with Vision-Language Reasoning

    Yu Wang, Yi Wang, Rui Dai +4

    cs.CVcs.AIcs.CYarXiv:2601.10477v22026
  19. Enhancing Underwater Imagery using Generative Adversarial Networks

    Cameron Fabbri, Md Jahidul Islam, Junaed Sattar

    cs.CVcs.ROarXiv:1801.04011v12018
  20. nnFormer: Interleaved Transformer for Volumetric Segmentation

    Hong-Yu Zhou, Jiansen Guo, Yinghao Zhang +3

    cs.CVarXiv:2109.03201v62021
  21. VideoMaMa: Mask-Guided Video Matting via Generative Prior

    Sangbeom Lim, Seoung Wug Oh, Jiahui Huang +3

    cs.CVcs.AIarXiv:2601.14255v12026
  22. Gated Context Aggregation Network for Image Dehazing and Deraining

    Dongdong Chen, Mingming He, Qingnan Fan +5

    cs.CVarXiv:1811.08747v22018
  23. Structure and Content-Guided Video Synthesis with Diffusion Models

    Patrick Esser, Johnathan Chiu, Parmida Atighehchian +2

    cs.CVarXiv:2302.03011v12023
  24. A Comprehensive Overhaul of Feature Distillation

    Byeongho Heo, Jeesoo Kim, Sangdoo Yun +3

    cs.CVcs.LGarXiv:1904.01866v22019
  25. Learning a Deep Embedding Model for Zero-Shot Learning

    Li Zhang, Tao Xiang, Shaogang Gong

    cs.CVarXiv:1611.05088v42016
  26. GroupViT: Semantic Segmentation Emerges from Text Supervision

    Jiarui Xu, Shalini De Mello, Sifei Liu +4

    cs.CVarXiv:2202.11094v52022
  27. VIOLA: Towards Video In-Context Learning with Minimal Annotations

    Ryo Fujii, Hideo Saito, Ryo Hachiuma

    cs.CVcs.AIarXiv:2601.15549v12026
  28. GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection

    Shuguang Zhang, Junhong Lian, Guoxin Yu +2

    cs.CVcs.AIcs.CLarXiv:2601.20618v12026
  29. Deep Learning COVID-19 Features on CXR using Limited Training Data Sets

    Yujin Oh, Sangjoon Park, Jong Chul Ye

    eess.IVcs.CVcs.LGarXiv:2004.05758v22020
  30. Deep Pyramidal Residual Networks

    Dongyoon Han, Jiwhan Kim, Junmo Kim

    cs.CVarXiv:1610.02915v42016
  31. Visual Transformers: Token-based Image Representation and Processing for Computer Vision

    Bichen Wu, Chenfeng Xu, Xiaoliang Dai +7

    cs.CVcs.LGeess.IVarXiv:2006.03677v42020
  32. RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

    Hanze Dong, Wei Xiong, Deepanshu Goyal +7

    cs.LGcs.AIcs.CLarXiv:2304.06767v42023
  33. PLANING: A Loosely Coupled Triangle-Gaussian Framework for Streaming 3D Reconstruction

    Changjian Jiang, Kerui Ren, Xudong Li +8

    cs.CVarXiv:2601.22046v42026
  34. Medical Image Synthesis with Context-Aware Generative Adversarial Networks

    Dong Nie, Roger Trullo, Caroline Petitjean +2

    cs.CVarXiv:1612.05362v12016
  35. Multi-Modal Fusion Transformer for End-to-End Autonomous Driving

    Aditya Prakash, Kashyap Chitta, Andreas Geiger

    cs.CVcs.AIcs.LGarXiv:2104.09224v12021
  36. Variational Information Distillation for Knowledge Transfer

    Sungsoo Ahn, Shell Xu Hu, Andreas Damianou +2

    cs.CVcs.AIcs.LGarXiv:1904.05835v12019
  37. Glance and Focus Reinforcement for Pan-cancer Screening

    Linshan Wu, Jiaxin Zhuang, Hao Chen

    cs.CVarXiv:2601.19103v22026
  38. Big Self-Supervised Models Advance Medical Image Classification

    Shekoofeh Azizi, Basil Mustafa, Fiona Ryan +9

    eess.IVcs.CVcs.LGarXiv:2101.05224v22021
  39. VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

    Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah

    cs.CVarXiv:2608.22521v12026
  40. Revisiting Self-Supervised Visual Representation Learning

    Alexander Kolesnikov, Xiaohua Zhai, Lucas Beyer

    cs.CVarXiv:1901.09005v12019
  41. Conformer: Local Features Coupling Global Representations for Visual Recognition

    Zhiliang Peng, Wei Huang, Shanzhi Gu +4

    cs.CVarXiv:2105.03889v12021
  42. VISTA-PATH: An interactive foundation model for pathology image segmentation and quantitative analysis in computational pathology

    Peixian Liang, Songhao Li, Shunsuke Koga +5

    cs.CVarXiv:2601.16451v12026
  43. Rethinking Spatial Dimensions of Vision Transformers

    Byeongho Heo, Sangdoo Yun, Dongyoon Han +3

    cs.CVarXiv:2103.16302v22021
  44. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing

    Hui Zhang, Juntao Liu, Zongkai Liu +4

    cs.CVarXiv:2603.11593v22026
  45. Emu3: Next-Token Prediction is All You Need

    Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22

    cs.CVarXiv:2409.18869v12024
  46. SimpleGPT: Improving GPT via A Simple Normalization Strategy

    Marco Chen, Xianbiao Qi, Yelin He +2

    cs.LGcs.CLcs.CVarXiv:2602.01212v12026
  47. A Survey of Appearance Models in Visual Object Tracking

    Xi Li, Weiming Hu, Chunhua Shen +3

    cs.CVarXiv:1303.4803v12013
  48. Dynamic Memory Networks for Visual and Textual Question Answering

    Caiming Xiong, Stephen Merity, Richard Socher

    cs.NEcs.CLcs.CVarXiv:1603.01417v12016
  49. Toward Real-World Single Image Super-Resolution: A New Benchmark and A New Model

    Jianrui Cai, Hui Zeng, Hongwei Yong +2

    cs.CVarXiv:1904.00523v12019
  50. Bridge Damage Detection from Low-Light UAV Imagery via Degradation-Aware Mixture-of-Experts Enhancement

    Hu Wang, Hongxu Pu, Zhiqi Hu +2

    cs.CVarXiv:2608.23136v12026
  51. VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text

    Hassan Akbari, Liangzhe Yuan, Rui Qian +4

    cs.CVcs.AIcs.LGarXiv:2104.11178v32021
  52. What makes for effective detection proposals?

    Jan Hosang, Rodrigo Benenson, Piotr Dollár +1

    cs.CVarXiv:1502.05082v32015
  53. Multi-scale Interactive Network for Salient Object Detection

    Youwei Pang, Xiaoqi Zhao, Lihe Zhang +1

    cs.CVarXiv:2007.09062v12020
  54. WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching

    Yihan Wang, Jia Deng

    cs.CVarXiv:2603.24836v32026
  55. Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

    Youliang Zhang, Zhengguang Zhou, Zhentao Yu +11

    cs.CVcs.AIcs.CLarXiv:2602.01538v12026
  56. Enhancing Multi-Image Understanding through Delimiter Token Scaling

    Minyoung Lee, Yeji Park, Dongjun Hwang +3

    cs.CVarXiv:2602.01984v22026
  57. SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation

    Yang Zou, Jongheon Jeong, Latha Pemula +2

    cs.CVarXiv:2207.14315v12022
  58. Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar

    Taqi Hamoda, Nuno Gracias

    cs.CVarXiv:2608.23479v12026
  59. BlendedMVS: A Large-scale Dataset for Generalized Multi-view Stereo Networks

    Yao Yao, Zixin Luo, Shiwei Li +5

    cs.CVarXiv:1911.10127v22019
  60. What Remains Normal? Clean Images Miss Useful Near-Defect Normal Patches for Anomaly Detection

    Joongwon Chae, Runming Wang, Peiwu Qin

    cs.CVarXiv:2608.23299v12026