Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,021 to 7,080 of 18,815
HA-CCN: Hierarchical Attention-based Crowd Counting Network
Vishwanath A. Sindagi, Vishal M. Patel
cs.CVarXiv:1907.10255v12019Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Xiaokang Chen, Zhiyu Wu, Xingchao Liu +5
cs.AIcs.CLcs.CVarXiv:2501.17811v12025Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Moo Jin Kim, Chelsea Finn, Percy Liang
cs.ROcs.AIcs.CVarXiv:2502.19645v22025Neighborhood Contrastive Learning for Novel Class Discovery
Zhun Zhong, Enrico Fini, Subhankar Roy +3
cs.CVcs.AIcs.LGarXiv:2106.10731v12021Depth Anything 3: Recovering the Visual Space from Any Views
Haotong Lin, Sili Chen, Junhao Liew +5
cs.CVarXiv:2511.10647v12025Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Yang Yue, Zhiqi Chen, Rui Lu +5
cs.AIcs.CLcs.CVarXiv:2504.13837v52025LayoutVAE: Stochastic Scene Layout Generation From a Label Set
Akash Abdu Jyothi, Thibaut Durand, Jiawei He +2
cs.CVarXiv:1907.10719v32019Lightweight Probabilistic Deep Networks
Jochen Gast, Stefan Roth
cs.CVcs.LGstat.MLarXiv:1805.11327v12018Stochastic Optimization of Tree Tensor Networks
Marius Willner, Maximilian Scharf, André Uschmajew +2
math.OCcs.CVphysics.comp-pharXiv:2609.00870v12026SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Michael Tschannen, Alexey Gritsenko, Xiao Wang +11
cs.CVcs.AIarXiv:2502.14786v12025UI-VISA: U-Net Initialized Vascular Image Segmentation Architecture
Asees Kaur, Suzanne S. Sindi, Erica M. Rutter
cs.CVarXiv:2609.01598v12026Benchmarking Robustness of 3D Object Detection to Common Corruptions in Autonomous Driving
Yinpeng Dong, Caixin Kang, Jinlai Zhang +6
cs.CVcs.AIcs.CRarXiv:2303.11040v12023T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks
Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai
cs.CVarXiv:1808.01454v12018CQF-HMR: Continuous Quaternion Flows for Probabilistic 3D Human Mesh Recovery from a Single Image
Cuong Le, Bao-Long Tran, Pavlo Melnyk +3
cs.CVarXiv:2609.00995v12026MAC: Mining Activity Concepts for Language-based Temporal Localization
Runzhou Ge, Jiyang Gao, Kan Chen +1
cs.CVarXiv:1811.08925v12018MAST: A Memory-Augmented Self-supervised Tracker
Zihang Lai, Erika Lu, Weidi Xie
cs.CVcs.LGarXiv:2002.07793v22020Multi-View Transformer for 3D Visual Grounding
Shijia Huang, Yilun Chen, Jiaya Jia +1
cs.CVcs.CLarXiv:2204.02174v12022Deep Metric Learning for Few-Shot Image Classification: A Review of Recent Developments
Xiaoxu Li, Xiaochen Yang, Zhanyu Ma +1
cs.CVcs.LGarXiv:2105.08149v22021An Artificial Agent for Robust Image Registration
Rui Liao, Shun Miao, Pierre de Tournemire +4
cs.CVarXiv:1611.10336v12016COVID-19 Infection Localization and Severity Grading from Chest X-ray Images
Anas M. Tahir, Muhammad E. H. Chowdhury, Amith Khandakar +11
eess.IVcs.CVarXiv:2103.07985v12021RL-GAN-Net: A Reinforcement Learning Agent Controlled GAN Network for Real-Time Point Cloud Shape Completion
Muhammad Sarmad, Hyunjoo Jenny Lee, Young Min Kim
cs.CVcs.AIarXiv:1904.12304v12019Sample and Computation Redistribution for Efficient Face Detection
Jia Guo, Jiankang Deng, Alexandros Lattas +1
cs.CVarXiv:2105.04714v12021Balanced MSE for Imbalanced Visual Regression
Jiawei Ren, Mingyuan Zhang, Cunjun Yu +1
cs.CVarXiv:2203.16427v12022Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures
Can Polat, Mustafa Kurban, Erchin Serpedin +1
cs.CVcond-mat.mtrl-sciphysics.chem-pharXiv:2609.00663v12026Contrastive Audio-Visual Masked Autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu +4
cs.MMcs.CVcs.SDarXiv:2210.07839v42022Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks
Wei Xiong, Wenhan Luo, Lin Ma +2
cs.CVarXiv:1709.07592v320173D Registration with Maximal Cliques
Xiyu Zhang, Jiaqi Yang, Shikun Zhang +1
cs.CVarXiv:2305.10854v12023Temporal Recurrent Networks for Online Action Detection
Mingze Xu, Mingfei Gao, Yi-Ting Chen +2
cs.CVarXiv:1811.07391v22018One Prompt Is Enough: Watermark Laundering Through Foundation Image Models
Jidong Yang, Qi Li, Wei Zong +5
cs.CVcs.AIcs.CRarXiv:2609.01249v12026SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation
Robin Chan, Krzysztof Lis, Svenja Uhlemeyer +6
cs.CVarXiv:2104.14812v22021LDConv: Linear deformable convolution for improving convolutional neural networks
Xin Zhang, Yingze Song, Tingting Song +4
cs.CVarXiv:2311.11587v32023Perceptual Video Quality Assessment: A Survey
Xiongkuo Min, Huiyu Duan, Wei Sun +2
cs.MMcs.CVeess.IVarXiv:2402.03413v12024ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
Sethuraman T, Savya Khosla, Onkar Kishor Susladkar +6
cs.CVarXiv:2609.00505v12026TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models
Chao Zhou, Yiling Chen, Qi Chu +3
cs.CVcs.AIcs.MMarXiv:2609.01277v12026InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Jinguo Zhu, Weiyun Wang, Zhe Chen +48
cs.CVarXiv:2504.10479v32025VGGT: Visual Geometry Grounded Transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev +3
cs.CVarXiv:2503.11651v12025Summaries:한국어Learning Attributes Equals Multi-Source Domain Generalization
Chuang Gan, Tianbao Yang, Boqing Gong
cs.CVarXiv:1605.00743v12016Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
cs.CVcs.AIarXiv:2511.21631v22025BVI-DVC: A Training Database for Deep Video Compression
Di Ma, Fan Zhang, David R. Bull
eess.IVcs.CVarXiv:2003.13552v22020Qwen-Image Technical Report
Chenfei Wu, Jiahao Li, Jingren Zhou +36
cs.CVarXiv:2508.02324v12025Linking Points With Labels in 3D: A Review of Point Cloud Semantic Segmentation
Yuxing Xie, Jiaojiao Tian, Xiao Xiang Zhu
cs.CVcs.LGeess.IVarXiv:1908.08854v32019Fi-ImageNet-1k: An OOD Benchmark From the Inside of the ImageNet-1k Validation Set
Ruslan Rozumnyi, Matěj Suchánek, Tomáš Vojíř +2
cs.CVarXiv:2609.01027v12026InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3
cs.CLcs.CVcs.HCarXiv:2609.01383v12026InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Weiyun Wang, Zhangwei Gao, Lixin Gu +72
cs.CVarXiv:2508.18265v22025Fast Sampling of Diffusion Models via Operator Learning
Hongkai Zheng, Weili Nie, Arash Vahdat +2
cs.LGcs.CVarXiv:2211.13449v32022Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models
Jiale Xu, Xintao Wang, Weihao Cheng +4
cs.CVarXiv:2212.14704v22022DINOv3
Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23
cs.CVcs.LGarXiv:2508.10104v12025YOLOv12: Attention-Centric Real-Time Object Detectors
Yunjie Tian, Qixiang Ye, David Doermann
cs.CVcs.AIarXiv:2502.12524v12025SphereReID: Deep Hypersphere Manifold Embedding for Person Re-Identification
Xing Fan, Wei Jiang, Hao Luo +1
cs.CVarXiv:1807.00537v12018MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro +2
cs.IRcs.AIcs.CLarXiv:2609.01316v12026Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan, Ang Wang, Baole Ai +59
cs.CVarXiv:2503.20314v22025Qwen2.5-VL Technical Report
Shuai Bai, Keqin Chen, Xuejing Liu +24
cs.CVcs.CLarXiv:2502.13923v12025Predicting Risk of Developing Diabetic Retinopathy using Deep Learning
Ashish Bora, Siva Balasubramanian, Boris Babenko +13
eess.IVcs.CVarXiv:2008.04370v12020NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
Shengming Yin, Chenfei Wu, Huan Yang +13
cs.CVcs.AIarXiv:2303.12346v12023UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs
Yanwu Xu, Yang Zhao, Zhisheng Xiao +1
cs.CVarXiv:2311.09257v52023PolyTransform: Deep Polygon Transformer for Instance Segmentation
Justin Liang, Namdar Homayounfar, Wei-Chiu Ma +3
cs.CVarXiv:1912.02801v42019Joint Line Segmentation and Transcription for End-to-End Handwritten Paragraph Recognition
Théodore Bluche
cs.CVcs.LGcs.NEarXiv:1604.08352v12016HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical Images
Meirui Jiang, Zirui Wang, Qi Dou
eess.IVcs.AIcs.CVarXiv:2112.10775v32021InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring
Zhihao Yuan, Xu Yan, Yinghong Liao +4
cs.CVarXiv:2103.01128v22021Naturalistic Driver Intention and Path Prediction using Recurrent Neural Networks
Alex Zyner, Stewart Worrall, Eduardo Nebot
cs.CVarXiv:1807.09995v12018