Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,301 to 6,360 of 18,866

  1. Skeleton-Based Action Recognition with Multi-Stream Adaptive Graph Convolutional Networks

    Lei Shi, Yifan Zhang, Jian Cheng +1

    cs.CVarXiv:1912.06971v12019
  2. General $E(2)$-Equivariant Steerable CNNs

    Maurice Weiler, Gabriele Cesa

    cs.CVcs.LGeess.IVarXiv:1911.08251v22019
  3. Grape detection, segmentation and tracking using deep neural networks and three-dimensional association

    Thiago T. Santos, Leonardo L. de Souza, Andreza A. dos Santos +1

    cs.CVarXiv:1907.11819v32019
  4. EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

    Rui Yang, Hanyang Chen, Junyu Zhang +10

    cs.AIcs.CLcs.CVarXiv:2502.09560v32025
  5. VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

    Dominic Maggio, Hyungtae Lim, Luca Carlone

    cs.CVarXiv:2505.12549v22025
  6. ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

    Yongkang Li, Kaixin Xiong, Xiangyu Guo +12

    cs.CVcs.ROarXiv:2506.08052v22025
  7. In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

    Zechuan Zhang, Ji Xie, Yu Lu +2

    cs.CVarXiv:2504.20690v32025
  8. Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation

    Yawei Luo, Ping Liu, Tao Guan +2

    cs.CVcs.AIcs.LGarXiv:1904.00876v12019
  9. PiP: Planning-informed Trajectory Prediction for Autonomous Driving

    Haoran Song, Wenchao Ding, Yuxuan Chen +3

    cs.CVcs.ROarXiv:2003.11476v22020
  10. A new Backdoor Attack in CNNs by training set corruption without label poisoning

    Mauro Barni, Kassem Kallas, Benedetta Tondi

    cs.CRcs.CVcs.LGarXiv:1902.11237v12019
  11. FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review

    Ahmad Shawahna, Sadiq M. Sait, Aiman El-Maleh

    cs.NEcs.ARcs.CVarXiv:1901.00121v12019
  12. Emu3.5: Native Multimodal Models are World Learners

    Yufeng Cui, Honghao Chen, Haoge Deng +20

    cs.CVarXiv:2510.26583v12025
  13. Full Flow: Optical Flow Estimation By Global Optimization over Regular Grids

    Qifeng Chen, Vladlen Koltun

    cs.CVarXiv:1604.03513v12016
  14. Whole-Slide Mitosis Detection in H&E Breast Histology Using PHH3 as a Reference to Train Distilled Stain-Invariant Convolutional Networks

    David Tellez, Maschenka Balkenhol, Irene Otte-Holler +10

    cs.CVarXiv:1808.05896v12018
  15. Deep Clustering for Unsupervised Learning of Visual Features

    Mathilde Caron, Piotr Bojanowski, Armand Joulin +1

    cs.CVarXiv:1807.05520v22018
  16. Continuous Learning in Single-Incremental-Task Scenarios

    Davide Maltoni, Vincenzo Lomonaco

    cs.LGcs.AIcs.CVarXiv:1806.08568v32018
  17. Robust Registration of Calcium Images by Learned Contrast Synthesis

    John A. Bogovic, Philipp Hanslovsky, Allan Wong +1

    cs.CVarXiv:1511.01154v12015
  18. MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

    Fanqing Meng, Lingxiao Du, Zongkai Liu +12

    cs.CVarXiv:2503.07365v22025
  19. SpatialTracker: Tracking Any 2D Pixels in 3D Space

    Yuxi Xiao, Qianqian Wang, Shangzhan Zhang +4

    cs.CVarXiv:2404.04319v12024
  20. Unsupervised Out-of-Distribution Detection by Maximum Classifier Discrepancy

    Qing Yu, Kiyoharu Aizawa

    cs.CVarXiv:1908.04951v12019
  21. MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

    Sihan Yang, Runsen Xu, Yiman Xie +10

    cs.CVcs.CLarXiv:2505.23764v32025
  22. THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling

    Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou +2

    cs.CVcs.ROarXiv:2110.06607v32021
  23. Object Detection in Videos with Tubelet Proposal Networks

    Kai Kang, Hongsheng Li, Tong Xiao +4

    cs.CVarXiv:1702.06355v22017
  24. Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

    Ariel Ephrat, Inbar Mosseri, Oran Lang +5

    cs.SDcs.CVeess.ASarXiv:1804.03619v22018
  25. Deep learning in radiology: an overview of the concepts and a survey of the state of the art

    Maciej A. Mazurowski, Mateusz Buda, Ashirbani Saha +1

    cs.CVcs.LGstat.AParXiv:1802.08717v12018
  26. Adversarial Examples: Attacks and Defenses for Deep Learning

    Xiaoyong Yuan, Pan He, Qile Zhu +1

    cs.LGcs.CRcs.CVarXiv:1712.07107v32017
  27. Wing Loss for Robust Facial Landmark Localisation with Convolutional Neural Networks

    Zhen-Hua Feng, Josef Kittler, Muhammad Awais +2

    cs.CVarXiv:1711.06753v52017
  28. Machine Learning for the Geosciences: Challenges and Opportunities

    Anuj Karpatne, Imme Ebert-Uphoff, Sai Ravela +2

    cs.LGcs.AIcs.CVarXiv:1711.04708v12017
  29. A Review of Convolutional Neural Networks for Inverse Problems in Imaging

    Michael T. McCann, Kyong Hwan Jin, Michael Unser

    eess.IVcs.CVarXiv:1710.04011v12017
  30. Machine learning \& artificial intelligence in the quantum domain

    Vedran Dunjko, Hans J. Briegel

    quant-phcs.AIcs.CVarXiv:1709.02779v12017
  31. A Brief Survey of Deep Reinforcement Learning

    Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage +1

    cs.LGcs.AIcs.CVarXiv:1708.05866v22017
  32. ToxTrac: a fast and robust software for tracking organisms

    Alvaro Rodriquez, Hanqing Zhang, Jonatan Klaminder +3

    cs.CVarXiv:1706.02577v12017
  33. Deep Learning Microscopy

    Yair Rivenson, Zoltan Gorocs, Harun Gunaydin +3

    cs.LGcs.CVphysics.opticsarXiv:1705.04709v12017
  34. MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

    Jiazhen Pan, Che Liu, Junde Wu +6

    cs.CVcs.AIarXiv:2502.19634v22025
  35. Building Deep Networks on Grassmann Manifolds

    Zhiwu Huang, Jiqing Wu, Luc Van Gool

    cs.CVarXiv:1611.05742v32016
  36. MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

    Junzhe Li, Yutao Cui, Tao Huang +8

    cs.AIcs.CVarXiv:2507.21802v72025
  37. Multi-Label Classification with Label Graph Superimposing

    Ya Wang, Dongliang He, Fu Li +4

    cs.CVarXiv:1911.09243v12019
  38. RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

    Isaac Robinson, Peter Robicheaux, Matvei Popov +2

    cs.CVarXiv:2511.09554v22025
  39. Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking

    Pengyu Zhang, Jie Zhao, Dong Wang +2

    cs.CVarXiv:2007.02041v12020
  40. VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

    Chaoyou Fu, Haojia Lin, Xiong Wang +13

    cs.CVcs.SDeess.ASarXiv:2501.01957v42025
  41. Comparison of machine learning methods for classifying mediastinal lymph node metastasis of non-small cell lung cancer from 18F-FDG PET/CT images

    Hongkai Wang, Zongwei Zhou, Yingci Li +5

    cs.CVphysics.med-pharXiv:1702.02223v12017
  42. Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

    Haocheng Xi, Shuo Yang, Yilong Zhao +11

    cs.CVcs.LGarXiv:2502.01776v22025
  43. CNN-based Segmentation of Medical Imaging Data

    Baris Kayalibay, Grady Jensen, Patrick van der Smagt

    cs.CVarXiv:1701.03056v22017
  44. Prototypical Cross-domain Self-supervised Learning for Few-shot Unsupervised Domain Adaptation

    Xiangyu Yue, Zangwei Zheng, Shanghang Zhang +4

    cs.CVarXiv:2103.16765v12021
  45. ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

    Haoyu Fu, Diankun Zhang, Zongchuang Zhao +7

    cs.CVarXiv:2503.19755v12025
  46. Guiding Instruction-based Image Editing via Multimodal Large Language Models

    Tsu-Jui Fu, Wenze Hu, Xianzhi Du +3

    cs.CVarXiv:2309.17102v22023
  47. Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment

    Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller +2

    cs.CVarXiv:1612.01697v22016
  48. Predicting Human Eye Fixations via an LSTM-based Saliency Attentive Model

    Marcella Cornia, Lorenzo Baraldi, Giuseppe Serra +1

    cs.CVarXiv:1611.09571v42016
  49. Deep Convolutional Neural Network for Inverse Problems in Imaging

    Kyong Hwan Jin, Michael T. McCann, Emmanuel Froustey +1

    cs.CVarXiv:1611.03679v12016
  50. Deep image mining for diabetic retinopathy screening

    Gwenolé Quellec, Katia Charrière, Yassine Boudi +2

    cs.CVarXiv:1610.07086v32016
  51. Exploring Nearest Neighbor Approaches for Image Captioning

    Jacob Devlin, Saurabh Gupta, Ross Girshick +2

    cs.CVarXiv:1505.04467v12015
  52. Semi-Supervised Sparse Representation Based Classification for Face Recognition with Insufficient Labeled Samples

    Yuan Gao, Jiayi Ma, Alan L. Yuille

    cs.CVarXiv:1609.03279v22016
  53. Diffusion Based Unpaired Data Learning for Inverse Problems

    Chenglong Bao, Yiming Dang, Chenguang Duan +2

    cs.CVarXiv:2609.01370v12026
  54. Fish Disease Detection Using Image Based Machine Learning Technique in Aquaculture

    Md Shoaib Ahmed, Tanjim Taharat Aurpa, Md. Abul Kalam Azad

    cs.CVcs.LGarXiv:2105.03934v12021
  55. Scene Text Detection via Holistic, Multi-Channel Prediction

    Cong Yao, Xiang Bai, Nong Sang +3

    cs.CVarXiv:1606.09002v22016
  56. Diversified Visual Attention Networks for Fine-Grained Object Classification

    Bo Zhao, Xiao Wu, Jiashi Feng +2

    cs.CVarXiv:1606.08572v22016
  57. Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via Discretisation

    Stephen James, Kentaro Wada, Tristan Laidlow +1

    cs.ROcs.AIcs.CVarXiv:2106.12534v22021
  58. Context-aware Deep Feature Compression for High-speed Visual Tracking

    Jongwon Choi, Hyung Jin Chang, Tobias Fischer +5

    cs.CVarXiv:1803.10537v12018
  59. Jo-SRC: A Contrastive Approach for Combating Noisy Labels

    Yazhou Yao, Zeren Sun, Chuanyi Zhang +4

    cs.CVarXiv:2103.13029v12021
  60. Multi-Modal Masked Autoencoders for Medical Vision-and-Language Pre-Training

    Zhihong Chen, Yuhao Du, Jinpeng Hu +4

    cs.CVcs.CLarXiv:2209.07098v12022