Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,301 to 6,360 of 18,866
Skeleton-Based Action Recognition with Multi-Stream Adaptive Graph Convolutional Networks
Lei Shi, Yifan Zhang, Jian Cheng +1
cs.CVarXiv:1912.06971v12019General $E(2)$-Equivariant Steerable CNNs
Maurice Weiler, Gabriele Cesa
cs.CVcs.LGeess.IVarXiv:1911.08251v22019Grape detection, segmentation and tracking using deep neural networks and three-dimensional association
Thiago T. Santos, Leonardo L. de Souza, Andreza A. dos Santos +1
cs.CVarXiv:1907.11819v32019EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Rui Yang, Hanyang Chen, Junyu Zhang +10
cs.AIcs.CLcs.CVarXiv:2502.09560v32025VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
Dominic Maggio, Hyungtae Lim, Luca Carlone
cs.CVarXiv:2505.12549v22025ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
Yongkang Li, Kaixin Xiong, Xiangyu Guo +12
cs.CVcs.ROarXiv:2506.08052v22025In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Zechuan Zhang, Ji Xie, Yu Lu +2
cs.CVarXiv:2504.20690v32025Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
Yawei Luo, Ping Liu, Tao Guan +2
cs.CVcs.AIcs.LGarXiv:1904.00876v12019PiP: Planning-informed Trajectory Prediction for Autonomous Driving
Haoran Song, Wenchao Ding, Yuxuan Chen +3
cs.CVcs.ROarXiv:2003.11476v22020A new Backdoor Attack in CNNs by training set corruption without label poisoning
Mauro Barni, Kassem Kallas, Benedetta Tondi
cs.CRcs.CVcs.LGarXiv:1902.11237v12019FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
Ahmad Shawahna, Sadiq M. Sait, Aiman El-Maleh
cs.NEcs.ARcs.CVarXiv:1901.00121v12019Emu3.5: Native Multimodal Models are World Learners
Yufeng Cui, Honghao Chen, Haoge Deng +20
cs.CVarXiv:2510.26583v12025Full Flow: Optical Flow Estimation By Global Optimization over Regular Grids
Qifeng Chen, Vladlen Koltun
cs.CVarXiv:1604.03513v12016Whole-Slide Mitosis Detection in H&E Breast Histology Using PHH3 as a Reference to Train Distilled Stain-Invariant Convolutional Networks
David Tellez, Maschenka Balkenhol, Irene Otte-Holler +10
cs.CVarXiv:1808.05896v12018Deep Clustering for Unsupervised Learning of Visual Features
Mathilde Caron, Piotr Bojanowski, Armand Joulin +1
cs.CVarXiv:1807.05520v22018Continuous Learning in Single-Incremental-Task Scenarios
Davide Maltoni, Vincenzo Lomonaco
cs.LGcs.AIcs.CVarXiv:1806.08568v32018Robust Registration of Calcium Images by Learned Contrast Synthesis
John A. Bogovic, Philipp Hanslovsky, Allan Wong +1
cs.CVarXiv:1511.01154v12015MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Fanqing Meng, Lingxiao Du, Zongkai Liu +12
cs.CVarXiv:2503.07365v22025SpatialTracker: Tracking Any 2D Pixels in 3D Space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang +4
cs.CVarXiv:2404.04319v12024Unsupervised Out-of-Distribution Detection by Maximum Classifier Discrepancy
Qing Yu, Kiyoharu Aizawa
cs.CVarXiv:1908.04951v12019MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Sihan Yang, Runsen Xu, Yiman Xie +10
cs.CVcs.CLarXiv:2505.23764v32025THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling
Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou +2
cs.CVcs.ROarXiv:2110.06607v32021Object Detection in Videos with Tubelet Proposal Networks
Kai Kang, Hongsheng Li, Tong Xiao +4
cs.CVarXiv:1702.06355v22017Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation
Ariel Ephrat, Inbar Mosseri, Oran Lang +5
cs.SDcs.CVeess.ASarXiv:1804.03619v22018Deep learning in radiology: an overview of the concepts and a survey of the state of the art
Maciej A. Mazurowski, Mateusz Buda, Ashirbani Saha +1
cs.CVcs.LGstat.AParXiv:1802.08717v12018Adversarial Examples: Attacks and Defenses for Deep Learning
Xiaoyong Yuan, Pan He, Qile Zhu +1
cs.LGcs.CRcs.CVarXiv:1712.07107v32017Wing Loss for Robust Facial Landmark Localisation with Convolutional Neural Networks
Zhen-Hua Feng, Josef Kittler, Muhammad Awais +2
cs.CVarXiv:1711.06753v52017Machine Learning for the Geosciences: Challenges and Opportunities
Anuj Karpatne, Imme Ebert-Uphoff, Sai Ravela +2
cs.LGcs.AIcs.CVarXiv:1711.04708v12017A Review of Convolutional Neural Networks for Inverse Problems in Imaging
Michael T. McCann, Kyong Hwan Jin, Michael Unser
eess.IVcs.CVarXiv:1710.04011v12017Machine learning \& artificial intelligence in the quantum domain
Vedran Dunjko, Hans J. Briegel
quant-phcs.AIcs.CVarXiv:1709.02779v12017A Brief Survey of Deep Reinforcement Learning
Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage +1
cs.LGcs.AIcs.CVarXiv:1708.05866v22017ToxTrac: a fast and robust software for tracking organisms
Alvaro Rodriquez, Hanqing Zhang, Jonatan Klaminder +3
cs.CVarXiv:1706.02577v12017Deep Learning Microscopy
Yair Rivenson, Zoltan Gorocs, Harun Gunaydin +3
cs.LGcs.CVphysics.opticsarXiv:1705.04709v12017MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
Jiazhen Pan, Che Liu, Junde Wu +6
cs.CVcs.AIarXiv:2502.19634v22025Building Deep Networks on Grassmann Manifolds
Zhiwu Huang, Jiqing Wu, Luc Van Gool
cs.CVarXiv:1611.05742v32016MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Junzhe Li, Yutao Cui, Tao Huang +8
cs.AIcs.CVarXiv:2507.21802v72025Multi-Label Classification with Label Graph Superimposing
Ya Wang, Dongliang He, Fu Li +4
cs.CVarXiv:1911.09243v12019RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
Isaac Robinson, Peter Robicheaux, Matvei Popov +2
cs.CVarXiv:2511.09554v22025Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking
Pengyu Zhang, Jie Zhao, Dong Wang +2
cs.CVarXiv:2007.02041v12020VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Chaoyou Fu, Haojia Lin, Xiong Wang +13
cs.CVcs.SDeess.ASarXiv:2501.01957v42025Comparison of machine learning methods for classifying mediastinal lymph node metastasis of non-small cell lung cancer from 18F-FDG PET/CT images
Hongkai Wang, Zongwei Zhou, Yingci Li +5
cs.CVphysics.med-pharXiv:1702.02223v12017Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
Haocheng Xi, Shuo Yang, Yilong Zhao +11
cs.CVcs.LGarXiv:2502.01776v22025CNN-based Segmentation of Medical Imaging Data
Baris Kayalibay, Grady Jensen, Patrick van der Smagt
cs.CVarXiv:1701.03056v22017Prototypical Cross-domain Self-supervised Learning for Few-shot Unsupervised Domain Adaptation
Xiangyu Yue, Zangwei Zheng, Shanghang Zhang +4
cs.CVarXiv:2103.16765v12021ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Haoyu Fu, Diankun Zhang, Zongchuang Zhao +7
cs.CVarXiv:2503.19755v12025Guiding Instruction-based Image Editing via Multimodal Large Language Models
Tsu-Jui Fu, Wenze Hu, Xianzhi Du +3
cs.CVarXiv:2309.17102v22023Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment
Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller +2
cs.CVarXiv:1612.01697v22016Predicting Human Eye Fixations via an LSTM-based Saliency Attentive Model
Marcella Cornia, Lorenzo Baraldi, Giuseppe Serra +1
cs.CVarXiv:1611.09571v42016Deep Convolutional Neural Network for Inverse Problems in Imaging
Kyong Hwan Jin, Michael T. McCann, Emmanuel Froustey +1
cs.CVarXiv:1611.03679v12016Deep image mining for diabetic retinopathy screening
Gwenolé Quellec, Katia Charrière, Yassine Boudi +2
cs.CVarXiv:1610.07086v32016Exploring Nearest Neighbor Approaches for Image Captioning
Jacob Devlin, Saurabh Gupta, Ross Girshick +2
cs.CVarXiv:1505.04467v12015Semi-Supervised Sparse Representation Based Classification for Face Recognition with Insufficient Labeled Samples
Yuan Gao, Jiayi Ma, Alan L. Yuille
cs.CVarXiv:1609.03279v22016Diffusion Based Unpaired Data Learning for Inverse Problems
Chenglong Bao, Yiming Dang, Chenguang Duan +2
cs.CVarXiv:2609.01370v12026Fish Disease Detection Using Image Based Machine Learning Technique in Aquaculture
Md Shoaib Ahmed, Tanjim Taharat Aurpa, Md. Abul Kalam Azad
cs.CVcs.LGarXiv:2105.03934v12021Scene Text Detection via Holistic, Multi-Channel Prediction
Cong Yao, Xiang Bai, Nong Sang +3
cs.CVarXiv:1606.09002v22016Diversified Visual Attention Networks for Fine-Grained Object Classification
Bo Zhao, Xiao Wu, Jiashi Feng +2
cs.CVarXiv:1606.08572v22016Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation
Stephen James, Kentaro Wada, Tristan Laidlow +1
cs.ROcs.AIcs.CVarXiv:2106.12534v22021Context-aware Deep Feature Compression for High-speed Visual Tracking
Jongwon Choi, Hyung Jin Chang, Tobias Fischer +5
cs.CVarXiv:1803.10537v12018Jo-SRC: A Contrastive Approach for Combating Noisy Labels
Yazhou Yao, Zeren Sun, Chuanyi Zhang +4
cs.CVarXiv:2103.13029v12021Multi-Modal Masked Autoencoders for Medical Vision-and-Language Pre-Training
Zhihong Chen, Yuhao Du, Jinpeng Hu +4
cs.CVcs.CLarXiv:2209.07098v12022