Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,841 to 15,900 of 18,839

  1. SpectralGPT: Spectral Remote Sensing Foundation Model

    Danfeng Hong, Bing Zhang, Xuyang Li +11

    cs.CVarXiv:2311.07113v32023
  2. ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst

    Mayank Bansal, Alex Krizhevsky, Abhijit Ogale

    cs.ROcs.CVcs.LGarXiv:1812.03079v12018
  3. HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

    Shangxuan Tian, Yanhui Chen, Carlos Queiroz

    cs.AIcs.CVcs.IRarXiv:2608.21792v12026
  4. iFSQ: Improving FSQ for Image Generation with 1 Line of Code

    Bin Lin, Zongjian Li, Yuwei Niu +9

    cs.CVarXiv:2601.17124v22026
  5. Contrastive Learning for Compact Single Image Dehazing

    Haiyan Wu, Yanyun Qu, Shaohui Lin +5

    cs.CVcs.AIarXiv:2104.09367v12021
  6. ClipCap: CLIP Prefix for Image Captioning

    Ron Mokady, Amir Hertz, Amit H. Bermano

    cs.CVarXiv:2111.09734v12021
  7. VidEoMT: Your ViT is Secretly Also a Video Segmentation Model

    Narges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero +4

    cs.CVarXiv:2602.17807v32026
  8. What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

    Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh +1

    cs.AIcs.CEcs.CVarXiv:2608.21761v12026
  9. Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

    Jiacheng Hua, Yishu Yin, Yuhang Wu +3

    cs.CVcs.CLarXiv:2603.23404v22026
  10. 360Anything: Geometry-Free Lifting of Images and Videos to 360°

    Ziyi Wu, Daniel Watson, Andrea Tagliasacchi +3

    cs.CVarXiv:2601.16192v22026
  11. Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields

    Jonathan T. Barron, Ben Mildenhall, Dor Verbin +2

    cs.CVcs.GRcs.LGarXiv:2304.06706v32023
  12. LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

    Yupan Huang, Tengchao Lv, Lei Cui +2

    cs.CLcs.CVarXiv:2204.08387v32022
  13. CondConv: Conditionally Parameterized Convolutions for Efficient Inference

    Brandon Yang, Gabriel Bender, Quoc V. Le +1

    cs.CVcs.AIcs.LGarXiv:1904.04971v32019
  14. Demystifying Action Space Design for Robotic Manipulation Policies

    Yuchun Feng, Jinliang Zheng, Zhihao Wang +5

    cs.ROcs.CVarXiv:2602.23408v22026
  15. Cardiologist-Level Arrhythmia Detection with Convolutional Neural Networks

    Pranav Rajpurkar, Awni Y. Hannun, Masoumeh Haghpanahi +2

    cs.CVarXiv:1707.01836v12017
  16. ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks

    Xiaohan Ding, Yuchen Guo, Guiguang Ding +1

    cs.CVcs.LGcs.NEarXiv:1908.03930v32019
  17. Test-Time Training with KV Binding Is Secretly Linear Attention

    Junchen Liu, Sven Elflein, Or Litany +2

    cs.LGcs.AIcs.CVarXiv:2602.21204v42026
  18. Not Just a Black Box: Learning Important Features Through Propagating Activation Differences

    Avanti Shrikumar, Peyton Greenside, Anna Shcherbina +1

    cs.LGcs.CVcs.NEarXiv:1605.01713v32016
  19. Gated Fusion Network for Single Image Dehazing

    Wenqi Ren, Lin Ma, Jiawei Zhang +4

    cs.CVarXiv:1804.00213v12018
  20. Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans

    Sida Peng, Yuanqing Zhang, Yinghao Xu +4

    cs.CVarXiv:2012.15838v22020
  21. HPatches: A benchmark and evaluation of handcrafted and learned local descriptors

    Vassileios Balntas, Karel Lenc, Andrea Vedaldi +1

    cs.CVarXiv:1704.05939v12017
  22. Perceiver IO: A General Architecture for Structured Inputs & Outputs

    Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac +12

    cs.LGcs.CLcs.CVarXiv:2107.14795v32021
  23. GeoWorld: Geometric World Models

    Zeyu Zhang, Danning Li, Ian Reid +1

    cs.CVcs.ROarXiv:2602.23058v22026
  24. Real time Detection of Lane Markers in Urban Streets

    Mohamed Aly

    cs.CVcs.ROarXiv:1411.7113v12014
  25. Self-labelling via simultaneous clustering and representation learning

    Yuki Markus Asano, Christian Rupprecht, Andrea Vedaldi

    cs.CVcs.NEarXiv:1911.05371v32019
  26. Objaverse-XL: A Universe of 10M+ 3D Objects

    Matt Deitke, Ruoshi Liu, Matthew Wallingford +14

    cs.CVcs.AIarXiv:2307.05663v12023
  27. FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space

    FSVideo Team, Qingyu Chen, Zhiyuan Fang +17

    cs.CVarXiv:2602.02092v12026
  28. Joint Unsupervised Learning of Deep Representations and Image Clusters

    Jianwei Yang, Devi Parikh, Dhruv Batra

    cs.CVcs.LGarXiv:1604.03628v32016
  29. MOT20: A benchmark for multi object tracking in crowded scenes

    Patrick Dendorfer, Hamid Rezatofighi, Anton Milan +6

    cs.CVarXiv:2003.09003v12020
  30. Pay Attention to MLPs

    Hanxiao Liu, Zihang Dai, David R. So +1

    cs.LGcs.CLcs.CVarXiv:2105.08050v22021
  31. MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

    Baorui Ma, Jiahui Yang, Donglin Di +5

    cs.CVcs.AIarXiv:2601.22054v22026
  32. YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation

    Ranjan Sapkota, Manoj Karkee

    cs.CVarXiv:2602.00168v12026
  33. Principal Neighbourhood Aggregation for Graph Nets

    Gabriele Corso, Luca Cavalleri, Dominique Beaini +2

    cs.LGcs.CVstat.MLarXiv:2004.05718v52020
  34. Representation Alignment for Just Image Transformers is not Easier than You Think

    Jaeyo Shin, Jiwook Kim, Hyunjung Shim

    cs.CVcs.LGarXiv:2603.14366v12026
  35. BlazePose: On-device Real-time Body Pose tracking

    Valentin Bazarevsky, Ivan Grishchenko, Karthik Raveendran +3

    cs.CVarXiv:2006.10204v12020
  36. Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

    David Bamman, Kent K. Chang, Allison Cooper +7

    cs.AIcs.CLcs.CVarXiv:2608.21430v12026
  37. Out-of-Distribution Detection with Deep Nearest Neighbors

    Yiyou Sun, Yifei Ming, Xiaojin Zhu +1

    cs.LGcs.CVarXiv:2204.06507v32022
  38. Learning Spatial Fusion for Single-Shot Object Detection

    Songtao Liu, Di Huang, Yunhong Wang

    cs.CVarXiv:1911.09516v22019
  39. Toward Cognitive Supersensing in Multimodal Large Language Model

    Boyi Li, Yifan Shen, Yuanzhe Liu +12

    cs.CVcs.AIarXiv:2602.01541v12026
  40. Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention

    Dvir Samuel, Issar Tzachor, Matan Levy +3

    cs.CVcs.AIarXiv:2602.01801v22026
  41. Unified Deep Supervised Domain Adaptation and Generalization

    Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh +1

    cs.CVarXiv:1709.10190v12017
  42. Point-E: A System for Generating 3D Point Clouds from Complex Prompts

    Alex Nichol, Heewoo Jun, Prafulla Dhariwal +2

    cs.CVcs.LGarXiv:2212.08751v12022
  43. Show, Don't Tell: Morphing Latent Reasoning into Image Generation

    Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu +6

    cs.CVarXiv:2602.02227v12026
  44. CHAOS Challenge -- Combined (CT-MR) Healthy Abdominal Organ Segmentation

    A. Emre Kavur, N. Sinem Gezer, Mustafa Barış +24

    eess.IVcs.CVarXiv:2001.06535v32020
  45. Towards Deep Neural Network Architectures Robust to Adversarial Examples

    Shixiang Gu, Luca Rigazio

    cs.LGcs.CVcs.NEarXiv:1412.5068v42014
  46. VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

    Qing'an Liu, Juntong Feng, Yuhao Wang +6

    cs.CVarXiv:2602.04802v32026
  47. Olaf-World: Orienting Latent Actions for Video World Modeling

    Yuxin Jiang, Yuchao Gu, Ivor W. Tsang +1

    cs.CVcs.AIcs.LGarXiv:2602.10104v22026
  48. Deep-COVID: Predicting COVID-19 From Chest X-Ray Images Using Deep Transfer Learning

    Shervin Minaee, Rahele Kafieh, Milan Sonka +2

    cs.CVarXiv:2004.09363v32020
  49. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation

    Tianwei Xiong, Jun Hao Liew, Zilong Huang +3

    cs.CVarXiv:2603.12267v12026
  50. GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning

    GigaBrain Team, Boyuan Wang, Bohan Li +23

    cs.CVarXiv:2602.12099v22026
  51. Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

    Weicai Yan, Yuhong Dai, Qi Ran +6

    cs.CVarXiv:2603.03447v42026
  52. SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

    Xinyao Zhang, Wenkai Dong, Yuxin Song +10

    cs.CVarXiv:2603.19228v12026
  53. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

    Xiaoshi Wu, Yiming Hao, Keqiang Sun +4

    cs.CVcs.AIcs.DBarXiv:2306.09341v22023
  54. Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set

    Yu Deng, Jiaolong Yang, Sicheng Xu +3

    cs.CVarXiv:1903.08527v22019
  55. Mip-Splatting: Alias-free 3D Gaussian Splatting

    Zehao Yu, Anpei Chen, Binbin Huang +2

    cs.CVarXiv:2311.16493v12023
  56. LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset

    Royden Wagner, Omer Sahin Tas, Jaime Villa +18

    cs.CVcs.ROarXiv:2603.23607v22026
  57. Real-Time Seamless Single Shot 6D Object Pose Prediction

    Bugra Tekin, Sudipta N. Sinha, Pascal Fua

    cs.CVarXiv:1711.08848v52017
  58. Detecting Twenty-thousand Classes using Image-level Supervision

    Xingyi Zhou, Rohit Girdhar, Armand Joulin +2

    cs.CVarXiv:2201.02605v32022
  59. BitDance: Scaling Autoregressive Generative Models with Binary Tokens

    Yuang Ai, Jiaming Han, Shaobin Zhuang +8

    cs.CVcs.AIarXiv:2602.14041v22026
  60. More Images, More Problems? A Controlled Analysis of VLM Failure Modes

    Anurag Das, Adrian Bulat, Alberto Baldrati +4

    cs.CVarXiv:2601.07812v12026