Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
4,141 to 4,200 of 18,786
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch +8
cs.CVcs.MMcs.SDarXiv:1901.01342v22019Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
Zhengyao Fang, Pengyuan Lyu, Chengquan Zhang +3
cs.CVarXiv:2603.09480v22026Deformable Kernel Networks for Joint Image Filtering
Beomjun Kim, Jean Ponce, Bumsub Ham
cs.CVarXiv:1910.08373v32019Object Concepts Emerge from Motion
Boshi Li, Xiaohui Wang, Xiaoyang Wu +3
cs.CVarXiv:2609.04348v12026AnyText: Multilingual Visual Text Generation And Editing
Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He +2
cs.CVarXiv:2311.03054v52023STyMo: Fast and Controllable Few-Shot Motion Style Transfer
Jose Luis Ponton, Alexander Winkler, Ladislav Kavan +2
cs.GRcs.CVarXiv:2609.04500v12026Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
Leyang Li, Shilin Lu, Yan Ren +1
cs.CVcs.AIcs.CRarXiv:2504.12782v12025Machine Learning-Aided Operations and Communications of Unmanned Aerial Vehicles: A Contemporary Survey
Harrison Kurunathan, Hailong Huang, Kai Li +2
cs.ROcs.CVcs.LGarXiv:2211.04324v12022MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
Ke Wang, Junting Pan, Linda Wei +8
cs.CVcs.AIcs.CLarXiv:2505.10557v12025DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
Ivan Kapelyukh, Vitalis Vosylius, Edward Johns
cs.ROcs.CVcs.LGarXiv:2210.02438v32022Knowledge Transfer with Jacobian Matching
Suraj Srinivas, Francois Fleuret
cs.LGcs.CVarXiv:1803.00443v12018Multi-Compound Transformer for Accurate Biomedical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Huijie Wang +4
cs.CVarXiv:2106.14385v12021Real-Time Apple Detection System Using Embedded Systems With Hardware Accelerators: An Edge AI Application
Vittorio Mazzia, Francesco Salvetti, Aleem Khaliq +1
cs.CVeess.IVarXiv:2004.13410v12020Enhanced Spatio-Temporal Interaction Learning for Video Deraining: A Faster and Better Framework
Kaihao Zhang, Dongxu Li, Wenhan Luo +2
cs.CVarXiv:2103.12318v22021EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
Siyuan Huang, Liliang Chen, Pengfei Zhou +8
cs.ROcs.CVcs.LGarXiv:2501.01895v32025Detecting Unexpected Obstacles for Self-Driving Cars: Fusing Deep Learning and Geometric Modeling
Sebastian Ramos, Stefan Gehrig, Peter Pinggera +2
cs.CVcs.ROarXiv:1612.06573v12016DARD: Zero-Shot Degradation-Aware Retinex-Guided Diffusion for Low-Light Image Enhancement
Wenjie Cai, Yuezhe Yang, Jianyang Xia +2
cs.CVarXiv:2608.29243v12026Physics-Based Generative Adversarial Models for Image Restoration and Beyond
Jinshan Pan, Jiangxin Dong, Yang Liu +5
cs.CVarXiv:1808.00605v22018Automatic Detection of Blue-White Veil and Related Structures in Dermoscopy Images
M. Emre Celebi, Hitoshi Iyatomi, William V. Stoecker +4
cs.CVarXiv:1009.1013v12010Samba: Semantic Segmentation of Remotely Sensed Images with State Space Model
Qinfeng Zhu, Yuanzhi Cai, Yuan Fang +4
cs.CVarXiv:2404.01705v22024Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Tao Xie, Peishan Yang, Yudong Jin +8
cs.CVarXiv:2604.08542v12026Computational Depth Measurement in Thermographic Video: Overcoming Spatial Overfitting via Spatio-Temporal Decoupling
Zain Ul Abidin, Habeeban Memon, Junaid Ahmed
cs.AIcs.CVarXiv:2608.29223v12026GSPotential: Camera Potential Field for Sparse-View 3D Gaussian Splatting
Zeyuan An, Yanghang Xiao, Zhiying Leng +2
cs.CVarXiv:2608.29346v12026Recent Advances in Zero-shot Recognition
Yanwei Fu, Tao Xiang, Yu-Gang Jiang +3
cs.CVcs.AIcs.LGarXiv:1710.04837v12017Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis
Jie Li, Xingchen Zou, Yuxuan Liang
cs.CVarXiv:2608.29233v12026OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang Huang, Fangchen Liu, Letian Fu +5
cs.ROcs.CVarXiv:2503.03734v42025Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
Bernard Koch, Emily Denton, Alex Hanna +1
cs.LGcs.CLcs.CVarXiv:2112.01716v12021Deep Visual Domain Adaptation
Gabriela Csurka
cs.CVarXiv:2012.14176v12020INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval
Zhiwei Chen, Yupeng Hu, Zhiheng Fu +4
cs.CVarXiv:2604.18051v12026Tiny Machine Learning: Progress and Futures
Ji Lin, Ligeng Zhu, Wei-Ming Chen +2
cs.LGcs.AIcs.CVarXiv:2403.19076v22024Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion
Xu Lin, Ke Wang, Hui Kang +1
cs.CVcs.AIcs.LGarXiv:2609.04690v12026Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical Imaging
Rui Yan, Liangqiong Qu, Qingyue Wei +5
cs.CVcs.LGarXiv:2205.08576v22022UNIC: Unified In-Context Video Editing
Zixuan Ye, Xuanhua He, Quande Liu +7
cs.CVarXiv:2506.04216v12025HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning
Shiming Chen, Guo-Sen Xie, Yang Liu +5
cs.CVcs.LGarXiv:2109.15163v22021Collaborative Diffusion for Multi-Modal Face Generation and Editing
Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang +1
cs.CVarXiv:2304.10530v12023Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing
Xinghe Fu, Zhiyuan Yan, Taiping Yao +2
cs.CVarXiv:2501.04376v12025REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory
Ziniu Hu, Ahmet Iscen, Chen Sun +6
cs.CVcs.AIarXiv:2212.05221v22022Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
Yue Ma, Kunyu Feng, Xinhua Zhang +7
cs.CVarXiv:2506.04590v12025EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
Sheng Zhou, Junbin Xiao, Qingyun Li +6
cs.CVcs.MMarXiv:2502.07411v22025Learning Depth from Single Images with Deep Neural Network Embedding Focal Length
Lei He, Guanghui Wang, Zhanyi Hu
cs.CVarXiv:1803.10039v12018Diffusion-SDF: Conditional Generative Modeling of Signed Distance Functions
Gene Chou, Yuval Bahat, Felix Heide
cs.CVarXiv:2211.13757v22022Instruction Distillation: Text Instructions as Visual Examples
Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury
cs.CVarXiv:2608.28696v12026Adaptive Rectangular Convolution for Remote Sensing Pansharpening
Xueyang Wang, Zhixin Zheng, Jiandong Shao +2
cs.CVeess.IVarXiv:2503.00467v12025An efficient iterative thresholding method for image segmentation
Dong Wang, Haohan Li, Xiaoyu Wei +1
cs.CVmath.NAarXiv:1608.01431v22016Active Convolution: Learning the Shape of Convolution for Image Classification
Yunho Jeon, Junmo Kim
cs.CVarXiv:1703.09076v12017Learning Visual Importance for Graphic Designs and Data Visualizations
Zoya Bylinskii, Nam Wook Kim, Peter O'Donovan +6
cs.HCcs.CVarXiv:1708.02660v12017The Unusual Effectiveness of Averaging in GAN Training
Yasin Yazıcı, Chuan-Sheng Foo, Stefan Winkler +3
stat.MLcs.CVcs.LGarXiv:1806.04498v22018AvatarMe: Realistically Renderable 3D Facial Reconstruction "in-the-wild"
Alexandros Lattas, Stylianos Moschoglou, Baris Gecer +4
cs.CVcs.GRarXiv:2003.13845v12020Distilling Multi-modal Large Language Models for Autonomous Driving
Deepti Hegde, Rajeev Yasarla, Hong Cai +7
cs.CVcs.ROarXiv:2501.09757v12025Automatic adaptation of object detectors to new domains using self-training
Aruni RoyChowdhury, Prithvijit Chakrabarty, Ashish Singh +4
cs.CVcs.LGarXiv:1904.07305v12019VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids
Katja Schwarz, Axel Sauer, Michael Niemeyer +2
cs.CVarXiv:2206.07695v32022RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction
Zifan Wang, Ziang Ren, Pengyang Shi +7
cs.ROcs.CVarXiv:2608.28693v120263D Point Cloud Denoising using Graph Laplacian Regularization of a Low Dimensional Manifold Model
Jin Zeng, Gene Cheung, Michael Ng +2
cs.CVarXiv:1803.07252v22018TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models
Bangwei Guo, Xujiang Zhao, Yanchi Liu +7
cs.CVarXiv:2608.28701v12026Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Yue Yang, Ajay Patel, Matt Deitke +8
cs.CVcs.CLarXiv:2502.14846v22025LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation
Weibin Liao, Yinghao Zhu, Xinyuan Wang +3
eess.IVcs.CVarXiv:2403.05246v22024Fusion of Detected Objects in Text for Visual Question Answering
Chris Alberti, Jeffrey Ling, Michael Collins +1
cs.CLcs.CVcs.LGarXiv:1908.05054v22019Wide Compression: Tensor Ring Nets
Wenqi Wang, Yifan Sun, Brian Eriksson +2
cs.LGcs.CVstat.MLarXiv:1802.09052v12018Video-Based Palm-Vein Authentication under Challenging Conditions
Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh +3
cs.CVarXiv:2609.02776v12026A Survey on (M)LLM-Based GUI Agents
Fei Tang, Haolei Xu, Hang Zhang +12
cs.HCcs.AIcs.CLarXiv:2504.13865v22025