Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,961 to 16,020 of 18,822

  1. Stable Velocity: A Variance Perspective on Flow Matching

    Donglin Yang, Yongxing Zhang, Xin Yu +5

    cs.CVarXiv:2602.05435v22026
  2. Pseudo Numerical Methods for Diffusion Models on Manifolds

    Luping Liu, Yi Ren, Zhijie Lin +1

    cs.CVcs.LGmath.NAarXiv:2202.09778v22022
  3. Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge

    Oriol Vinyals, Alexander Toshev, Samy Bengio +1

    cs.CVarXiv:1609.06647v12016
  4. VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

    Zhongwei Ren, Yunchao Wei, Xiao Yu +5

    cs.CVarXiv:2602.10102v12026
  5. A White Paper on Neural Network Quantization

    Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3

    cs.LGcs.AIcs.CVarXiv:2106.08295v12021
  6. Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering

    Tao Lu, Mulin Yu, Linning Xu +4

    cs.CVarXiv:2312.00109v12023
  7. Hypergraph Convolution and Hypergraph Attention

    Song Bai, Feihu Zhang, Philip H. S. Torr

    cs.LGcs.CVstat.MLarXiv:1901.08150v22019
  8. Scaling Vision Transformers to 22 Billion Parameters

    Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

    cs.CVcs.AIcs.LGarXiv:2302.05442v12023
  9. Classification of COVID-19 in chest X-ray images using DeTraC deep convolutional neural network

    Asmaa Abbas, Mohammed M. Abdelsamea, Mohamed Medhat Gaber

    eess.IVcs.CVcs.LGarXiv:2003.13815v32020
  10. OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams

    Yibin Yan, Jilan Xu, Shangzhe Di +2

    cs.CVarXiv:2603.12265v12026
  11. The Creation and Detection of Deepfakes: A Survey

    Yisroel Mirsky, Wenke Lee

    cs.CVcs.LGeess.IVarXiv:2004.11138v32020
  12. Human Motion Trajectory Prediction: A Survey

    Andrey Rudenko, Luigi Palmieri, Michael Herman +3

    cs.ROcs.CVcs.LGarXiv:1905.06113v32019
  13. Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

    Chengwen Liu, Xiaomin Yu, Zhuoyue Chang +15

    cs.CVcs.AIarXiv:2601.06943v22026
  14. TextBoxes: A Fast Text Detector with a Single Deep Neural Network

    Minghui Liao, Baoguang Shi, Xiang Bai +2

    cs.CVarXiv:1611.06779v12016
  15. It Takes Two: A Duet of Periodicity and Directionality for Burst Flicker Removal

    Lishen Qu, Shihao Zhou, Jie Liang +3

    cs.CVarXiv:2603.22794v12026
  16. CityPersons: A Diverse Dataset for Pedestrian Detection

    Shanshan Zhang, Rodrigo Benenson, Bernt Schiele

    cs.CVarXiv:1702.05693v12017
  17. Bottom-up Object Detection by Grouping Extreme and Center Points

    Xingyi Zhou, Jiacheng Zhuo, Philipp Krähenbühl

    cs.CVarXiv:1901.08043v32019
  18. Generalizing to Unseen Domains via Adversarial Data Augmentation

    Riccardo Volpi, Hongseok Namkoong, Ozan Sener +3

    cs.CVarXiv:1805.12018v22018
  19. Feature Selective Anchor-Free Module for Single-Shot Object Detection

    Chenchen Zhu, Yihui He, Marios Savvides

    cs.CVarXiv:1903.00621v12019
  20. Learning a Convolutional Neural Network for Non-uniform Motion Blur Removal

    Jian Sun, Wenfei Cao, Zongben Xu +1

    cs.CVarXiv:1503.00593v32015
  21. Beyond Inferring Class Representatives: User-Level Privacy Leakage From Federated Learning

    Zhibo Wang, Mengkai Song, Zhifei Zhang +3

    cs.LGcs.CRcs.CVarXiv:1812.00535v32018
  22. The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics

    Xiangbo Gao, Mingyang Wu, Siyuan Yang +4

    cs.CVcs.AIarXiv:2603.14375v22026
  23. The MegaFace Benchmark: 1 Million Faces for Recognition at Scale

    Ira Kemelmacher-Shlizerman, Steve Seitz, Daniel Miller +1

    cs.CVarXiv:1512.00596v12015
  24. DeepVO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks

    Sen Wang, Ronald Clark, Hongkai Wen +1

    cs.CVcs.ROarXiv:1709.08429v12017
  25. Exploring Visual Relationship for Image Captioning

    Ting Yao, Yingwei Pan, Yehao Li +1

    cs.CVarXiv:1809.07041v12018
  26. Learning Deep Representations of Fine-grained Visual Descriptions

    Scott Reed, Zeynep Akata, Bernt Schiele +1

    cs.CVarXiv:1605.05395v12016
  27. Are We on the Right Way for Evaluating Large Vision-Language Models?

    Lin Chen, Jinsong Li, Xiaoyi Dong +8

    cs.CVarXiv:2403.20330v22024
  28. PointNetLK: Robust & Efficient Point Cloud Registration using PointNet

    Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan +1

    cs.CVarXiv:1903.05711v22019
  29. MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding

    Hejun Dong, Junbo Niu, Bin Wang +3

    cs.CVarXiv:2603.22458v12026
  30. WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

    Zhiyang Deng, Boran Zhang, Danze Chen +1

    cs.CVarXiv:2608.21439v12026
  31. DS-TransUNet:Dual Swin Transformer U-Net for Medical Image Segmentation

    Ailiang Lin, Bingzhi Chen, Jiayu Xu +2

    cs.CVarXiv:2106.06716v12021
  32. Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud

    Weijing Shi, Ragunathan, Rajkumar

    cs.CVarXiv:2003.01251v12020
  33. Long Range Arena: A Benchmark for Efficient Transformers

    Yi Tay, Mostafa Dehghani, Samira Abnar +7

    cs.LGcs.AIcs.CLarXiv:2011.04006v12020
  34. Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

    Xiaomin Yu, Yi Xin, Yuhui Zhang +12

    cs.CVcs.AIcs.MMarXiv:2602.07026v32026
  35. Appearance-Based Gaze Estimation in the Wild

    Xucong Zhang, Yusuke Sugano, Mario Fritz +1

    cs.CVarXiv:1504.02863v12015
  36. End-to-end Autonomous Driving: Challenges and Frontiers

    Li Chen, Penghao Wu, Kashyap Chitta +3

    cs.ROcs.AIcs.CVarXiv:2306.16927v32023
  37. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

    Sirui Xu, Samuel Schulter, Morteza Ziyadi +4

    cs.CVcs.GRcs.ROarXiv:2602.06035v12026
  38. Multimodal Chain-of-Thought Reasoning in Language Models

    Zhuosheng Zhang, Aston Zhang, Mu Li +3

    cs.CLcs.AIcs.CVarXiv:2302.00923v52023
  39. Genetic CNN

    Lingxi Xie, Alan Yuille

    cs.CVarXiv:1703.01513v12017
  40. Gliding vertex on the horizontal bounding box for multi-oriented object detection

    Yongchao Xu, Mingtao Fu, Qimeng Wang +4

    cs.CVarXiv:1911.09358v22019
  41. From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

    Xiangyan Qu, Zhenlong Yuan, Jing Tang +9

    cs.CVcs.AIcs.LGarXiv:2603.00141v32026
  42. UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience

    Zichuan Lin, Feiyu Liu, Yijun Yang +9

    cs.LGcs.AIcs.CVarXiv:2603.24533v12026
  43. Image Generation from Scene Graphs

    Justin Johnson, Agrim Gupta, Li Fei-Fei

    cs.CVcs.LGarXiv:1804.01622v12018
  44. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

    Jovana Kondic, Pengyuan Li, Dhiraj Joshi +24

    cs.CVcs.AIcs.CLarXiv:2603.27064v22026
  45. A Survey of Modern Deep Learning based Object Detection Models

    Syed Sahil Abbas Zaidi, Mohammad Samar Ansari, Asra Aslam +3

    cs.CVcs.LGeess.IVarXiv:2104.11892v22021
  46. Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks

    Wentao Zhu, Cuiling Lan, Junliang Xing +4

    cs.CVcs.LGarXiv:1603.07772v12016
  47. Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset

    Scott Ettinger, Shuyang Cheng, Benjamin Caine +15

    cs.CVcs.LGcs.ROarXiv:2104.10133v12021
  48. Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought

    Yunheng Li, Hangyi Kuang, Hengrui Zhang +4

    cs.CVarXiv:2603.22847v12026
  49. Deep Convolutional Neural Fields for Depth Estimation from a Single Image

    Fayao Liu, Chunhua Shen, Guosheng Lin

    cs.CVarXiv:1411.6387v22014
  50. Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

    Zichen Wen, Boxue Yang, Shuang Chen +31

    cs.CVcs.AIarXiv:2601.19325v12026
    Summaries:한국어
  51. Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP Framework

    Xu Ma, Can Qin, Haoxuan You +2

    cs.CVcs.AIarXiv:2202.07123v22022
  52. VisDA: The Visual Domain Adaptation Challenge

    Xingchao Peng, Ben Usman, Neela Kaushik +3

    cs.CVarXiv:1710.06924v22017
  53. VLS: Steering Pretrained Robot Policies via Vision-Language Models

    Shuo Liu, Ishneet Sukhvinder Singh, Yiqing Xu +2

    cs.ROcs.CVarXiv:2602.03973v12026
  54. Computer Vision for Autonomous Vehicles: Problems, Datasets and State of the Art

    Joel Janai, Fatma Güney, Aseem Behl +1

    cs.CVcs.ROarXiv:1704.05519v32017
  55. D2-Net: A Trainable CNN for Joint Detection and Description of Local Features

    Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla +4

    cs.CVarXiv:1905.03561v12019
  56. Deep Continuous Fusion for Multi-Sensor 3D Object Detection

    Ming Liang, Bin Yang, Shenlong Wang +1

    cs.CVarXiv:2012.10992v12020
  57. PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation

    Xin-Sheng Chen, Jiayu Zhu, Pei-lin Li +3

    cs.CVarXiv:2603.07244v12026
  58. Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

    Zengbin Wang, Xuecai Hu, Yong Wang +3

    cs.CVarXiv:2601.20354v22026
  59. RIVER: A Real-Time Interaction Benchmark for Video LLMs

    Yansong Shi, Qingsong Zhao, Tianxiang Jiang +3

    cs.CVarXiv:2603.03985v12026
  60. Dense Nested Attention Network for Infrared Small Target Detection

    Boyang Li, Chao Xiao, Longguang Wang +5

    cs.CVarXiv:2106.00487v32021