Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,141 to 4,200 of 18,786

  1. AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection

    Joseph Roth, Sourish Chaudhuri, Ondrej Klejch +8

    cs.CVcs.MMcs.SDarXiv:1901.01342v22019
  2. Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

    Zhengyao Fang, Pengyuan Lyu, Chengquan Zhang +3

    cs.CVarXiv:2603.09480v22026
  3. Deformable Kernel Networks for Joint Image Filtering

    Beomjun Kim, Jean Ponce, Bumsub Ham

    cs.CVarXiv:1910.08373v32019
  4. Object Concepts Emerge from Motion

    Boshi Li, Xiaohui Wang, Xiaoyang Wu +3

    cs.CVarXiv:2609.04348v12026
  5. AnyText: Multilingual Visual Text Generation And Editing

    Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He +2

    cs.CVarXiv:2311.03054v52023
  6. STyMo: Fast and Controllable Few-Shot Motion Style Transfer

    Jose Luis Ponton, Alexander Winkler, Ladislav Kavan +2

    cs.GRcs.CVarXiv:2609.04500v12026
  7. Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts

    Leyang Li, Shilin Lu, Yan Ren +1

    cs.CVcs.AIcs.CRarXiv:2504.12782v12025
  8. Machine Learning-Aided Operations and Communications of Unmanned Aerial Vehicles: A Contemporary Survey

    Harrison Kurunathan, Hailong Huang, Kai Li +2

    cs.ROcs.CVcs.LGarXiv:2211.04324v12022
  9. MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning

    Ke Wang, Junting Pan, Linda Wei +8

    cs.CVcs.AIcs.CLarXiv:2505.10557v12025
  10. DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics

    Ivan Kapelyukh, Vitalis Vosylius, Edward Johns

    cs.ROcs.CVcs.LGarXiv:2210.02438v32022
  11. Knowledge Transfer with Jacobian Matching

    Suraj Srinivas, Francois Fleuret

    cs.LGcs.CVarXiv:1803.00443v12018
  12. Multi-Compound Transformer for Accurate Biomedical Image Segmentation

    Yuanfeng Ji, Ruimao Zhang, Huijie Wang +4

    cs.CVarXiv:2106.14385v12021
  13. Real-Time Apple Detection System Using Embedded Systems With Hardware Accelerators: An Edge AI Application

    Vittorio Mazzia, Francesco Salvetti, Aleem Khaliq +1

    cs.CVeess.IVarXiv:2004.13410v12020
  14. Enhanced Spatio-Temporal Interaction Learning for Video Deraining: A Faster and Better Framework

    Kaihao Zhang, Dongxu Li, Wenhan Luo +2

    cs.CVarXiv:2103.12318v22021
  15. EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

    Siyuan Huang, Liliang Chen, Pengfei Zhou +8

    cs.ROcs.CVcs.LGarXiv:2501.01895v32025
  16. Detecting Unexpected Obstacles for Self-Driving Cars: Fusing Deep Learning and Geometric Modeling

    Sebastian Ramos, Stefan Gehrig, Peter Pinggera +2

    cs.CVcs.ROarXiv:1612.06573v12016
  17. DARD: Zero-Shot Degradation-Aware Retinex-Guided Diffusion for Low-Light Image Enhancement

    Wenjie Cai, Yuezhe Yang, Jianyang Xia +2

    cs.CVarXiv:2608.29243v12026
  18. Physics-Based Generative Adversarial Models for Image Restoration and Beyond

    Jinshan Pan, Jiangxin Dong, Yang Liu +5

    cs.CVarXiv:1808.00605v22018
  19. Automatic Detection of Blue-White Veil and Related Structures in Dermoscopy Images

    M. Emre Celebi, Hitoshi Iyatomi, William V. Stoecker +4

    cs.CVarXiv:1009.1013v12010
  20. Samba: Semantic Segmentation of Remotely Sensed Images with State Space Model

    Qinfeng Zhu, Yuanzhi Cai, Yuan Fang +4

    cs.CVarXiv:2404.01705v22024
  21. Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

    Tao Xie, Peishan Yang, Yudong Jin +8

    cs.CVarXiv:2604.08542v12026
  22. Computational Depth Measurement in Thermographic Video: Overcoming Spatial Overfitting via Spatio-Temporal Decoupling

    Zain Ul Abidin, Habeeban Memon, Junaid Ahmed

    cs.AIcs.CVarXiv:2608.29223v12026
  23. GSPotential: Camera Potential Field for Sparse-View 3D Gaussian Splatting

    Zeyuan An, Yanghang Xiao, Zhiying Leng +2

    cs.CVarXiv:2608.29346v12026
  24. Recent Advances in Zero-shot Recognition

    Yanwei Fu, Tao Xiang, Yu-Gang Jiang +3

    cs.CVcs.AIcs.LGarXiv:1710.04837v12017
  25. Generalization over Memorization: Generalization-Aware Diffusion Adaptation for Single-Image Multi-View Synthesis

    Jie Li, Xingchen Zou, Yuxuan Liang

    cs.CVarXiv:2608.29233v12026
  26. OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

    Huang Huang, Fangchen Liu, Letian Fu +5

    cs.ROcs.CVarXiv:2503.03734v42025
  27. Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

    Bernard Koch, Emily Denton, Alex Hanna +1

    cs.LGcs.CLcs.CVarXiv:2112.01716v12021
  28. Deep Visual Domain Adaptation

    Gabriela Csurka

    cs.CVarXiv:2012.14176v12020
  29. INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

    Zhiwei Chen, Yupeng Hu, Zhiheng Fu +4

    cs.CVarXiv:2604.18051v12026
  30. Tiny Machine Learning: Progress and Futures

    Ji Lin, Ligeng Zhu, Wei-Ming Chen +2

    cs.LGcs.AIcs.CVarXiv:2403.19076v22024
  31. Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

    Xu Lin, Ke Wang, Hui Kang +1

    cs.CVcs.AIcs.LGarXiv:2609.04690v12026
  32. Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical Imaging

    Rui Yan, Liangqiong Qu, Qingyue Wei +5

    cs.CVcs.LGarXiv:2205.08576v22022
  33. UNIC: Unified In-Context Video Editing

    Zixuan Ye, Xuanhua He, Quande Liu +7

    cs.CVarXiv:2506.04216v12025
  34. HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning

    Shiming Chen, Guo-Sen Xie, Yang Liu +5

    cs.CVcs.LGarXiv:2109.15163v22021
  35. Collaborative Diffusion for Multi-Modal Face Generation and Editing

    Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang +1

    cs.CVarXiv:2304.10530v12023
  36. Exploring Unbiased Deepfake Detection via Token-Level Shuffling and Mixing

    Xinghe Fu, Zhiyuan Yan, Taiping Yao +2

    cs.CVarXiv:2501.04376v12025
  37. REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory

    Ziniu Hu, Ahmet Iscen, Chen Sun +6

    cs.CVcs.AIarXiv:2212.05221v22022
  38. Follow-Your-Creation: Empowering 4D Creation through Video Inpainting

    Yue Ma, Kunyu Feng, Xinhua Zhang +7

    cs.CVarXiv:2506.04590v12025
  39. EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

    Sheng Zhou, Junbin Xiao, Qingyun Li +6

    cs.CVcs.MMarXiv:2502.07411v22025
  40. Learning Depth from Single Images with Deep Neural Network Embedding Focal Length

    Lei He, Guanghui Wang, Zhanyi Hu

    cs.CVarXiv:1803.10039v12018
  41. Diffusion-SDF: Conditional Generative Modeling of Signed Distance Functions

    Gene Chou, Yuval Bahat, Felix Heide

    cs.CVarXiv:2211.13757v22022
  42. Instruction Distillation: Text Instructions as Visual Examples

    Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury

    cs.CVarXiv:2608.28696v12026
  43. Adaptive Rectangular Convolution for Remote Sensing Pansharpening

    Xueyang Wang, Zhixin Zheng, Jiandong Shao +2

    cs.CVeess.IVarXiv:2503.00467v12025
  44. An efficient iterative thresholding method for image segmentation

    Dong Wang, Haohan Li, Xiaoyu Wei +1

    cs.CVmath.NAarXiv:1608.01431v22016
  45. Active Convolution: Learning the Shape of Convolution for Image Classification

    Yunho Jeon, Junmo Kim

    cs.CVarXiv:1703.09076v12017
  46. Learning Visual Importance for Graphic Designs and Data Visualizations

    Zoya Bylinskii, Nam Wook Kim, Peter O'Donovan +6

    cs.HCcs.CVarXiv:1708.02660v12017
  47. The Unusual Effectiveness of Averaging in GAN Training

    Yasin Yazıcı, Chuan-Sheng Foo, Stefan Winkler +3

    stat.MLcs.CVcs.LGarXiv:1806.04498v22018
  48. AvatarMe: Realistically Renderable 3D Facial Reconstruction "in-the-wild"

    Alexandros Lattas, Stylianos Moschoglou, Baris Gecer +4

    cs.CVcs.GRarXiv:2003.13845v12020
  49. Distilling Multi-modal Large Language Models for Autonomous Driving

    Deepti Hegde, Rajeev Yasarla, Hong Cai +7

    cs.CVcs.ROarXiv:2501.09757v12025
  50. Automatic adaptation of object detectors to new domains using self-training

    Aruni RoyChowdhury, Prithvijit Chakrabarty, Ashish Singh +4

    cs.CVcs.LGarXiv:1904.07305v12019
  51. VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel Grids

    Katja Schwarz, Axel Sauer, Michael Niemeyer +2

    cs.CVarXiv:2206.07695v32022
  52. RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

    Zifan Wang, Ziang Ren, Pengyang Shi +7

    cs.ROcs.CVarXiv:2608.28693v12026
  53. 3D Point Cloud Denoising using Graph Laplacian Regularization of a Low Dimensional Manifold Model

    Jin Zeng, Gene Cheung, Michael Ng +2

    cs.CVarXiv:1803.07252v22018
  54. TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models

    Bangwei Guo, Xujiang Zhao, Yanchi Liu +7

    cs.CVarXiv:2608.28701v12026
  55. Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

    Yue Yang, Ajay Patel, Matt Deitke +8

    cs.CVcs.CLarXiv:2502.14846v22025
  56. LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation

    Weibin Liao, Yinghao Zhu, Xinyuan Wang +3

    eess.IVcs.CVarXiv:2403.05246v22024
  57. Fusion of Detected Objects in Text for Visual Question Answering

    Chris Alberti, Jeffrey Ling, Michael Collins +1

    cs.CLcs.CVcs.LGarXiv:1908.05054v22019
  58. Wide Compression: Tensor Ring Nets

    Wenqi Wang, Yifan Sun, Brian Eriksson +2

    cs.LGcs.CVstat.MLarXiv:1802.09052v12018
  59. Video-Based Palm-Vein Authentication under Challenging Conditions

    Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh +3

    cs.CVarXiv:2609.02776v12026
  60. A Survey on (M)LLM-Based GUI Agents

    Fei Tang, Haolei Xu, Hang Zhang +12

    cs.HCcs.AIcs.CLarXiv:2504.13865v22025