Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,901 to 6,960 of 18,807
Designing Versatile Samples for Learned Trajectory Scoring
Yaguang Li, Jiaru Zhang, Chuheng Wei +2
cs.ROcs.CVarXiv:2609.01799v12026Deep Learning on Chest X-ray Images to Detect and Evaluate Pneumonia Cases at the Era of COVID-19
Karim Hammoudi, Halim Benhabiles, Mahmoud Melkemi +4
eess.IVcs.CVcs.LGarXiv:2004.03399v12020LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory
Kun-Yang Yu, Yingzhe Li, Hongyu Xu +8
cs.CVcs.ROarXiv:2609.02350v12026Deep convolutional neural networks for pedestrian detection
Denis Tomè, Federico Monti, Luca Baroffio +3
cs.CVarXiv:1510.03608v52015F*: An Interpretable Transformation of the F-measure
David J. Hand, Peter Christen, Nishadi Kirielle
cs.LGcs.AIcs.CVarXiv:2008.00103v32020Sim4CV: A Photo-Realistic Simulator for Computer Vision Applications
Matthias Müller, Vincent Casser, Jean Lahoud +2
cs.CVarXiv:1708.05869v22017Bag of Freebies for Training Object Detection Neural Networks
Zhi Zhang, Tong He, Hang Zhang +3
cs.CVarXiv:1902.04103v32019CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction
Menghao Li, Linjie Mu, Yin Wang +4
cs.CVarXiv:2609.02401v12026RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer
Jian Wang, Chenhui Gou, Qiman Wu +4
cs.CVarXiv:2210.07124v12022TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images
Shubham Paliwal, Vishwanath D, Rohit Rahul +2
cs.CVcs.LGeess.IVarXiv:2001.01469v12020Continual Learning with Pre-Trained Models: A Survey
Da-Wei Zhou, Hai-Long Sun, Jingyi Ning +2
cs.LGcs.CVarXiv:2401.16386v22024HookNet: multi-resolution convolutional neural networks for semantic segmentation in histopathology whole-slide images
Mart van Rijthoven, Maschenka Balkenhol, Karina Siliņa +2
eess.IVcs.CVcs.LGarXiv:2006.12230v12020TGANet: Text-guided attention for improved polyp segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci +1
eess.IVcs.CVcs.LGarXiv:2205.04280v12022DeepFake Detection: Current Challenges and Next Steps
Siwei Lyu
cs.CVarXiv:2003.09234v12020EdgeStereo: A Context Integrated Residual Pyramid Network for Stereo Matching
Xiao Song, Xu Zhao, Hanwen Hu +1
cs.CVarXiv:1803.05196v32018End-to-end Generative Pretraining for Multimodal Video Captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab +1
cs.CVcs.AIcs.CLarXiv:2201.08264v22022Transferable Contrastive Network for Generalized Zero-Shot Learning
Huajie Jiang, Ruiping Wang, Shiguang Shan +1
cs.CVarXiv:1908.05832v12019Arbitrary Style Transfer with Deep Feature Reshuffle
Shuyang Gu, Congliang Chen, Jing Liao +1
cs.CVarXiv:1805.04103v42018ResRep: Lossless CNN Pruning via Decoupling Remembering and Forgetting
Xiaohan Ding, Tianxiang Hao, Jianchao Tan +4
cs.LGcs.CVeess.IVarXiv:2007.03260v42020Polar Transformer Networks
Carlos Esteves, Christine Allen-Blanchette, Xiaowei Zhou +1
cs.CVarXiv:1709.01889v32017Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment
Zihao Wang, Xi Xiang, Yuwen Sun +5
cs.CVarXiv:2609.02573v120263D Shape Reconstruction from Sketches via Multi-view Convolutional Networks
Zhaoliang Lun, Matheus Gadelha, Evangelos Kalogerakis +2
cs.CVcs.GRarXiv:1707.06375v32017Copycat CNN: Stealing Knowledge by Persuading Confession with Random Non-Labeled Data
Jacson Rodrigues Correia-Silva, Rodrigo F. Berriel, Claudine Badue +2
cs.CVstat.MLarXiv:1806.05476v12018Knowledge Distillation via Route Constrained Optimization
Xiao Jin, Baoyun Peng, Yichao Wu +5
cs.LGcs.CVarXiv:1904.09149v12019In-Hand Object Rotation via Rapid Motor Adaptation
Haozhi Qi, Ashish Kumar, Roberto Calandra +2
cs.ROcs.AIcs.CVarXiv:2210.04887v12022Dynamic Multimodal Instance Segmentation guided by natural language queries
Edgar Margffoy-Tuay, Juan C. Pérez, Emilio Botero +1
cs.CVarXiv:1807.02257v22018Recognition of Ischaemia and Infection in Diabetic Foot Ulcers: Dataset and Techniques
Manu Goyal, Neil Reeves, Satyan Rajbhandari +3
eess.IVcs.CVarXiv:1908.05317v42019Refusion: Enabling Large-Size Realistic Image Restoration with Latent-Space Diffusion Models
Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2
cs.CVarXiv:2304.08291v12023AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo Network
Zizhuang Wei, Qingtian Zhu, Chen Min +2
cs.CVarXiv:2108.03824v12021Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization
Luke Melas-Kyriazi, Christian Rupprecht, Iro Laina +1
cs.CVcs.AIarXiv:2205.07839v12022Deep Learning Approaches on Image Captioning: A Review
Taraneh Ghandi, Hamidreza Pourreza, Hamidreza Mahyar
cs.CVarXiv:2201.12944v52022Inversion by Direct Iteration: An Alternative to Denoising Diffusion for Image Restoration
Mauricio Delbracio, Peyman Milanfar
eess.IVcs.CVcs.LGarXiv:2303.11435v52023SCOP: Scientific Control for Reliable Neural Network Pruning
Yehui Tang, Yunhe Wang, Yixing Xu +4
cs.CVcs.LGarXiv:2010.10732v22020Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks
Rachel Lea Draelos, Lawrence Carin
eess.IVcs.CVcs.LGarXiv:2011.08891v42020GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing
Lingxiao Li, Max Whitton, Ledell Wu +1
cs.CVarXiv:2609.00525v12026EarthLD: Towards Unified Open-World Landslide Understanding via Vision-Language Guided Diffusion Models
Yuanchao Su, Lianru Gao, Mengying Jiang +3
cs.CVarXiv:2609.00712v12026Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
Mohammadamin Barekatain, Miquel Martí, Hsueh-Fu Shih +4
cs.CVarXiv:1706.03038v22017Intriguing Properties of Contrastive Losses
Ting Chen, Calvin Luo, Lala Li
cs.LGcs.AIcs.CVarXiv:2011.02803v32020GSVA: Generalized Segmentation via Multimodal Large Language Models
Zhuofan Xia, Dongchen Han, Yizeng Han +3
cs.CVarXiv:2312.10103v32023MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Xiangxiang Chu, Limeng Qiao, Xinyu Zhang +8
cs.CVcs.AIarXiv:2402.03766v12024BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval
Gencer Sumbul, Arne de Wall, Tristan Kreuziger +6
cs.CVarXiv:2105.07921v22021Graph Convolution for Multimodal Information Extraction from Visually Rich Documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang +1
cs.IRcs.CVcs.LGarXiv:1903.11279v12019Step1X-Edit: A Practical Framework for General Image Editing
Shiyu Liu, Yucheng Han, Peng Xing +21
cs.CVarXiv:2504.17761v52025YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
Mengqi Lei, Siqi Li, Yihong Wu +7
cs.CVarXiv:2506.17733v22025Dynamics Based 3D Skeletal Hand Tracking
Stan Melax, Leonid Keselman, Sterling Orsten
cs.CVcs.GRarXiv:1705.07640v12017VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Boqiang Zhang, Kehan Li, Zesen Cheng +12
cs.CVarXiv:2501.13106v42025Residual Attention: A Simple but Effective Method for Multi-Label Recognition
Ke Zhu, Jianxin Wu
cs.CVarXiv:2108.02456v22021Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Tianyuan Yuan, Zibin Dong, Yicheng Liu +1
cs.CVcs.AIarXiv:2603.16666v22026DoFE: Domain-oriented Feature Embedding for Generalizable Fundus Image Segmentation on Unseen Datasets
Shujun Wang, Lequan Yu, Kang Li +3
cs.CVarXiv:2010.06208v12020MedGemma Technical Report
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri +78
cs.AIcs.CLcs.CVarXiv:2507.05201v42025Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Wenxuan Huang, Bohan Jia, Zijie Zhai +7
cs.CVcs.AIcs.CLarXiv:2503.06749v42025Grounded Human-Object Interaction Hotspots from Video
Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman
cs.CVarXiv:1812.04558v22018Comprehensive Graph-conditional Similarity Preserving Network for Unsupervised Cross-modal Hashing
Jun Yu, Hao Zhou, Yibing Zhan +1
cs.IRcs.CVarXiv:2012.13538v12020Understanding metric-related pitfalls in image analysis validation
Annika Reinke, Minu D. Tizabi, Michael Baumgartner +75
cs.CVarXiv:2302.01790v42023MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Yubo Ma, Yuhang Zang, Liangyu Chen +13
cs.CVcs.CLarXiv:2407.01523v32024Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
cs.CLcs.AIcs.CVarXiv:2509.17765v12025A Benchmark for Lidar Sensors in Fog: Is Detection Breaking Down?
Mario Bijelic, Tobias Gruber, Werner Ritter
cs.CVarXiv:1912.03251v12019Spot the conversation: speaker diarisation in the wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani +2
cs.SDcs.CVeess.ASarXiv:2007.01216v32020SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection
Rui Fan, Hengli Wang, Peide Cai +1
cs.CVcs.ROeess.IVarXiv:2008.11351v12020InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation
Vanshika Vats, Ashwani Rathee, James Davis
cs.CVcs.AIarXiv:2609.02002v12026