Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,841 to 15,900 of 18,839
SpectralGPT: Spectral Remote Sensing Foundation Model
Danfeng Hong, Bing Zhang, Xuyang Li +11
cs.CVarXiv:2311.07113v32023ChauffeurNet: Learning to Drive by Imitating the Best and Synthesizing the Worst
Mayank Bansal, Alex Krizhevsky, Abhijit Ogale
cs.ROcs.CVcs.LGarXiv:1812.03079v12018HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries
Shangxuan Tian, Yanhui Chen, Carlos Queiroz
cs.AIcs.CVcs.IRarXiv:2608.21792v12026iFSQ: Improving FSQ for Image Generation with 1 Line of Code
Bin Lin, Zongjian Li, Yuwei Niu +9
cs.CVarXiv:2601.17124v22026Contrastive Learning for Compact Single Image Dehazing
Haiyan Wu, Yanyun Qu, Shaohui Lin +5
cs.CVcs.AIarXiv:2104.09367v12021ClipCap: CLIP Prefix for Image Captioning
Ron Mokady, Amir Hertz, Amit H. Bermano
cs.CVarXiv:2111.09734v12021VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
Narges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero +4
cs.CVarXiv:2602.17807v32026What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh +1
cs.AIcs.CEcs.CVarXiv:2608.21761v12026Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
Jiacheng Hua, Yishu Yin, Yuhang Wu +3
cs.CVcs.CLarXiv:2603.23404v22026360Anything: Geometry-Free Lifting of Images and Videos to 360°
Ziyi Wu, Daniel Watson, Andrea Tagliasacchi +3
cs.CVarXiv:2601.16192v22026Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields
Jonathan T. Barron, Ben Mildenhall, Dor Verbin +2
cs.CVcs.GRcs.LGarXiv:2304.06706v32023LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Yupan Huang, Tengchao Lv, Lei Cui +2
cs.CLcs.CVarXiv:2204.08387v32022CondConv: Conditionally Parameterized Convolutions for Efficient Inference
Brandon Yang, Gabriel Bender, Quoc V. Le +1
cs.CVcs.AIcs.LGarXiv:1904.04971v32019Demystifying Action Space Design for Robotic Manipulation Policies
Yuchun Feng, Jinliang Zheng, Zhihao Wang +5
cs.ROcs.CVarXiv:2602.23408v22026Cardiologist-Level Arrhythmia Detection with Convolutional Neural Networks
Pranav Rajpurkar, Awni Y. Hannun, Masoumeh Haghpanahi +2
cs.CVarXiv:1707.01836v12017ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks
Xiaohan Ding, Yuchen Guo, Guiguang Ding +1
cs.CVcs.LGcs.NEarXiv:1908.03930v32019Test-Time Training with KV Binding Is Secretly Linear Attention
Junchen Liu, Sven Elflein, Or Litany +2
cs.LGcs.AIcs.CVarXiv:2602.21204v42026Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina +1
cs.LGcs.CVcs.NEarXiv:1605.01713v32016Gated Fusion Network for Single Image Dehazing
Wenqi Ren, Lin Ma, Jiawei Zhang +4
cs.CVarXiv:1804.00213v12018Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans
Sida Peng, Yuanqing Zhang, Yinghao Xu +4
cs.CVarXiv:2012.15838v22020HPatches: A benchmark and evaluation of handcrafted and learned local descriptors
Vassileios Balntas, Karel Lenc, Andrea Vedaldi +1
cs.CVarXiv:1704.05939v12017Perceiver IO: A General Architecture for Structured Inputs & Outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac +12
cs.LGcs.CLcs.CVarXiv:2107.14795v32021GeoWorld: Geometric World Models
Zeyu Zhang, Danning Li, Ian Reid +1
cs.CVcs.ROarXiv:2602.23058v22026Real time Detection of Lane Markers in Urban Streets
Mohamed Aly
cs.CVcs.ROarXiv:1411.7113v12014Self-labelling via simultaneous clustering and representation learning
Yuki Markus Asano, Christian Rupprecht, Andrea Vedaldi
cs.CVcs.NEarXiv:1911.05371v32019Objaverse-XL: A Universe of 10M+ 3D Objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford +14
cs.CVcs.AIarXiv:2307.05663v12023FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
FSVideo Team, Qingyu Chen, Zhiyuan Fang +17
cs.CVarXiv:2602.02092v12026Joint Unsupervised Learning of Deep Representations and Image Clusters
Jianwei Yang, Devi Parikh, Dhruv Batra
cs.CVcs.LGarXiv:1604.03628v32016MOT20: A benchmark for multi object tracking in crowded scenes
Patrick Dendorfer, Hamid Rezatofighi, Anton Milan +6
cs.CVarXiv:2003.09003v12020Pay Attention to MLPs
Hanxiao Liu, Zihang Dai, David R. So +1
cs.LGcs.CLcs.CVarXiv:2105.08050v22021MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
cs.CVcs.AIarXiv:2601.22054v22026YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation
Ranjan Sapkota, Manoj Karkee
cs.CVarXiv:2602.00168v12026Principal Neighbourhood Aggregation for Graph Nets
Gabriele Corso, Luca Cavalleri, Dominique Beaini +2
cs.LGcs.CVstat.MLarXiv:2004.05718v52020Representation Alignment for Just Image Transformers is not Easier than You Think
Jaeyo Shin, Jiwook Kim, Hyunjung Shim
cs.CVcs.LGarXiv:2603.14366v12026BlazePose: On-device Real-time Body Pose tracking
Valentin Bazarevsky, Ivan Grishchenko, Karthik Raveendran +3
cs.CVarXiv:2006.10204v12020Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
David Bamman, Kent K. Chang, Allison Cooper +7
cs.AIcs.CLcs.CVarXiv:2608.21430v12026Out-of-Distribution Detection with Deep Nearest Neighbors
Yiyou Sun, Yifei Ming, Xiaojin Zhu +1
cs.LGcs.CVarXiv:2204.06507v32022Learning Spatial Fusion for Single-Shot Object Detection
Songtao Liu, Di Huang, Yunhong Wang
cs.CVarXiv:1911.09516v22019Toward Cognitive Supersensing in Multimodal Large Language Model
Boyi Li, Yifan Shen, Yuanzhe Liu +12
cs.CVcs.AIarXiv:2602.01541v12026Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Dvir Samuel, Issar Tzachor, Matan Levy +3
cs.CVcs.AIarXiv:2602.01801v22026Unified Deep Supervised Domain Adaptation and Generalization
Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh +1
cs.CVarXiv:1709.10190v12017Point-E: A System for Generating 3D Point Clouds from Complex Prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal +2
cs.CVcs.LGarXiv:2212.08751v12022Show, Don't Tell: Morphing Latent Reasoning into Image Generation
Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu +6
cs.CVarXiv:2602.02227v12026CHAOS Challenge -- Combined (CT-MR) Healthy Abdominal Organ Segmentation
A. Emre Kavur, N. Sinem Gezer, Mustafa Barış +24
eess.IVcs.CVarXiv:2001.06535v32020Towards Deep Neural Network Architectures Robust to Adversarial Examples
Shixiang Gu, Luca Rigazio
cs.LGcs.CVcs.NEarXiv:1412.5068v42014VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
Qing'an Liu, Juntong Feng, Yuhao Wang +6
cs.CVarXiv:2602.04802v32026Olaf-World: Orienting Latent Actions for Video World Modeling
Yuxin Jiang, Yuchao Gu, Ivor W. Tsang +1
cs.CVcs.AIcs.LGarXiv:2602.10104v22026Deep-COVID: Predicting COVID-19 From Chest X-Ray Images Using Deep Transfer Learning
Shervin Minaee, Rahele Kafieh, Milan Sonka +2
cs.CVarXiv:2004.09363v32020EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
Tianwei Xiong, Jun Hao Liew, Zilong Huang +3
cs.CVarXiv:2603.12267v12026GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning
GigaBrain Team, Boyuan Wang, Bohan Li +23
cs.CVarXiv:2602.12099v22026Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
Weicai Yan, Yuhong Dai, Qi Ran +6
cs.CVarXiv:2603.03447v42026SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Xinyao Zhang, Wenkai Dong, Yuxin Song +10
cs.CVarXiv:2603.19228v12026Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Xiaoshi Wu, Yiming Hao, Keqiang Sun +4
cs.CVcs.AIcs.DBarXiv:2306.09341v22023Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set
Yu Deng, Jiaolong Yang, Sicheng Xu +3
cs.CVarXiv:1903.08527v22019Mip-Splatting: Alias-free 3D Gaussian Splatting
Zehao Yu, Anpei Chen, Binbin Huang +2
cs.CVarXiv:2311.16493v12023LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
Royden Wagner, Omer Sahin Tas, Jaime Villa +18
cs.CVcs.ROarXiv:2603.23607v22026Real-Time Seamless Single Shot 6D Object Pose Prediction
Bugra Tekin, Sudipta N. Sinha, Pascal Fua
cs.CVarXiv:1711.08848v52017Detecting Twenty-thousand Classes using Image-level Supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin +2
cs.CVarXiv:2201.02605v32022BitDance: Scaling Autoregressive Generative Models with Binary Tokens
Yuang Ai, Jiaming Han, Shaobin Zhuang +8
cs.CVcs.AIarXiv:2602.14041v22026More Images, More Problems? A Controlled Analysis of VLM Failure Modes
Anurag Das, Adrian Bulat, Alberto Baldrati +4
cs.CVarXiv:2601.07812v12026