Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

7,141 to 7,200 of 18,817

  1. What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

    Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3

    cs.CVarXiv:2609.01551v12026
  2. Extreme View Synthesis

    Inchang Choi, Orazio Gallo, Alejandro Troccoli +2

    cs.CVarXiv:1812.04777v22018
  3. PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices

    Xiaolong Ma, Fu-Ming Guo, Wei Niu +5

    cs.LGcs.CVcs.DCarXiv:1909.05073v42019
  4. Editable Visual Design

    Junyan Ye, Wei Liu, Dongzhi Jiang +9

    cs.CVcs.CLarXiv:2609.04034v12026
  5. Introduction to the Bag of Features Paradigm for Image Classification and Retrieval

    Stephen O'Hara, Bruce A. Draper

    cs.CVcs.IRarXiv:1101.3354v12011
  6. Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice

    Xiaojiang Peng, Limin Wang, Xingxing Wang +1

    cs.CVarXiv:1405.4506v12014
  7. Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

    Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta +2

    cs.CVcs.HCarXiv:2609.02512v12026
  8. AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates

    Ning Liu, Xiaolong Ma, Zhiyuan Xu +3

    cs.LGcs.AIcs.CVarXiv:1907.03141v22019
  9. Learning to Generate Images with Perceptual Similarity Metrics

    Jake Snell, Karl Ridgeway, Renjie Liao +3

    cs.LGcs.CVarXiv:1511.06409v32015
  10. Structured Prediction Helps 3D Human Motion Modelling

    Emre Aksan, Manuel Kaufmann, Otmar Hilliges

    cs.CVarXiv:1910.09070v12019
  11. CenterFormer: Center-based Transformer for 3D Object Detection

    Zixiang Zhou, Xiangchen Zhao, Yu Wang +2

    cs.CVarXiv:2209.05588v12022
  12. Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching

    Youmin Zhang, Yimin Chen, Xiao Bai +4

    cs.CVarXiv:1909.03751v22019
  13. Evaluating the Impact of Intensity Normalization on MR Image Synthesis

    Jacob C. Reinhold, Blake E. Dewey, Aaron Carass +1

    cs.CVarXiv:1812.04652v12018
  14. Real-time Driver Drowsiness Detection for Android Application Using Deep Neural Networks Techniques

    Rateb Jabbar, Khalifa Al-Khalifa, Mohamed Kharbeche +3

    cs.CVcs.HCarXiv:1811.01627v12018
  15. Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

    Shanji Liu, Kelu Yao, Junxiao Xue +5

    cs.CVarXiv:2609.02187v12026
  16. Human-centric Indoor Scene Synthesis Using Stochastic Grammar

    Siyuan Qi, Yixin Zhu, Siyuan Huang +2

    cs.CVarXiv:1808.08473v12018
  17. KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection

    Zhouyuan Huai, Meiqi Wan, Yan Yang +4

    cs.CVarXiv:2609.02077v12026
  18. Looking Beyond Appearances: Synthetic Training Data for Deep CNNs in Re-identification

    Igor Barros Barbosa, Marco Cristani, Barbara Caputo +2

    cs.CVarXiv:1701.03153v22017
  19. CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation

    Hao Sun, Hongyi Wang, Jiaqing Liu +2

    cs.MMcs.CLcs.CVarXiv:2207.14087v32022
  20. Low Frequency Adversarial Perturbation

    Chuan Guo, Jared S. Frank, Kilian Q. Weinberger

    cs.CVarXiv:1809.08758v22018
  21. Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement

    Yuhui Wu, Chen Pan, Guoqing Wang +4

    cs.CVarXiv:2304.07039v12023
  22. A Neural Temporal Model for Human Motion Prediction

    Anand Gopalakrishnan, Ankur Mali, Dan Kifer +2

    cs.CVarXiv:1809.03036v52018
  23. ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

    Jitai Hao, Ke Yang, Qiang Huang +1

    cs.CVcs.CLarXiv:2609.02780v12026
  24. Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and Adaptation

    Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin +1

    cs.CVarXiv:1909.09675v12019
  25. TACO: Trash Annotations in Context for Litter Detection

    Pedro F Proença, Pedro Simões

    cs.CVarXiv:2003.06975v22020
  26. CLIP-Driven Fine-grained Text-Image Person Re-identification

    Shuanglin Yan, Neng Dong, Liyan Zhang +1

    cs.CVarXiv:2210.10276v12022
  27. SpatialBot: Precise Spatial Understanding with Vision Language Models

    Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4

    cs.CVarXiv:2406.13642v72024
  28. ELEGANT: Exchanging Latent Encodings with GAN for Transferring Multiple Face Attributes

    Taihong Xiao, Jiapeng Hong, Jinwen Ma

    cs.CVarXiv:1803.10562v22018
  29. Pose Guided Structured Region Ensemble Network for Cascaded Hand Pose Estimation

    Xinghao Chen, Guijin Wang, Hengkai Guo +1

    cs.CVarXiv:1708.03416v22017
  30. Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

    Keen You, Haotian Zhang, Eldon Schoop +5

    cs.CVcs.CLcs.HCarXiv:2404.05719v12024
  31. SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation

    Qin Lu, Zheyang Jing, Yujie Yang +3

    cs.CVcs.AIarXiv:2609.02247v12026
  32. CASIA-SURF: A Large-scale Multi-modal Benchmark for Face Anti-spoofing

    Shifeng Zhang, Ajian Liu, Jun Wan +5

    cs.CVarXiv:1908.10654v22019
  33. Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation

    Xiaoxiao Li, Chen Change Loy

    cs.CVarXiv:1803.04242v22018
  34. Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models

    Haoning Wu, Zicheng Zhang, Erli Zhang +11

    cs.CVcs.MMarXiv:2311.06783v12023
  35. HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

    Zhaorun Chen, Zhuokai Zhao, Hongyin Luo +3

    cs.CVcs.AIcs.LGarXiv:2403.00425v22024
  36. Map-Guided Curriculum Domain Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation

    Christos Sakaridis, Dengxin Dai, Luc Van Gool

    cs.CVarXiv:2005.14553v22020
    Summaries:한국어
  37. VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM

    Sangmin Song, Sarath Kodagoda, Marc G. Carmichael +4

    cs.CVcs.AIarXiv:2609.00775v12026
  38. FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age

    Kimmo Kärkkäinen, Jungseock Joo

    cs.CVcs.LGarXiv:1908.04913v12019
  39. Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain

    Daizong Liu, Junhao Dong, Zhiyuan Ma +6

    cs.CVarXiv:2609.00788v12026
  40. Real-time Cardiovascular MR with Spatio-temporal Artifact Suppression using Deep Learning - Proof of Concept in Congenital Heart Disease

    Andreas Hauptmann, Simon Arridge, Felix Lucka +2

    cs.CVcs.NEarXiv:1803.05192v32018
  41. Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification

    Lin Wu, Yang Wang, Junbin Gao +1

    cs.CVarXiv:1808.01911v22018
  42. RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing

    Kaiyue Kang, Qixuan He, Peijin Wang +9

    cs.CVarXiv:2609.00814v12026
  43. Abnormality Detection and Localization in Chest X-Rays using Deep Convolutional Neural Networks

    Mohammad Tariqul Islam, Md Abdul Aowal, Ahmed Tahseen Minhaz +1

    cs.CVarXiv:1705.09850v32017
  44. Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison

    Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit

    cs.CVarXiv:2609.01530v12026
  45. HorizonNet: Learning Room Layout with 1D Representation and Pano Stretch Data Augmentation

    Cheng Sun, Chi-Wei Hsiao, Min Sun +1

    cs.CVarXiv:1901.03861v22019
  46. OVANet: One-vs-All Network for Universal Domain Adaptation

    Kuniaki Saito, Kate Saenko

    cs.CVarXiv:2104.03344v42021
  47. ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence

    Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao +1

    cs.CVarXiv:2609.01344v12026
  48. Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot Learning

    Zhiqiang Shen, Zechun Liu, Jie Qin +2

    cs.CVcs.AIcs.LGarXiv:2102.03983v12021
  49. Panoptic NeRF: 3D-to-2D Label Transfer for Panoptic Urban Scene Segmentation

    Xiao Fu, Shangzhan Zhang, Tianrun Chen +5

    cs.CVarXiv:2203.15224v22022
  50. TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

    Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9

    cs.CVarXiv:2609.01899v12026
  51. Fourier Space Losses for Efficient Perceptual Image Super-Resolution

    Dario Fuoli, Luc Van Gool, Radu Timofte

    eess.IVcs.CVarXiv:2106.00783v12021
  52. Deep Learning for Human Affect Recognition: Insights and New Developments

    Philipp V. Rouast, Marc T. P. Adam, Raymond Chiong

    cs.LGcs.AIcs.CVarXiv:1901.02884v12019
  53. A Survey of Deep Learning for Mathematical Reasoning

    Pan Lu, Liang Qiu, Wenhao Yu +2

    cs.AIcs.CLcs.CVarXiv:2212.10535v22022
  54. Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment

    Mohamad Jouni, Aurélien Godet, Mauro Dalla Mura

    eess.IVcs.CVarXiv:2609.01060v12026
  55. Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study

    Samuel Ritter, David G. T. Barrett, Adam Santoro +1

    stat.MLcs.CVcs.LGarXiv:1706.08606v22017
  56. UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations

    Chad Melton, PhD., Annabelle Kelton

    cs.CVcs.DBarXiv:2609.01738v12026
  57. Loss Aware Post-training Quantization

    Yury Nahshan, Brian Chmiel, Chaim Baskin +4

    cs.LGcs.CVarXiv:1911.07190v22019
  58. One-pass Multi-task Networks with Cross-task Guided Attention for Brain Tumor Segmentation

    Chenhong Zhou, Changxing Ding, Xinchao Wang +2

    cs.CVcs.AIcs.LGarXiv:1906.01796v22019
  59. Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

    Kang Liao, Yihang Luo, Xiao-Ming Wu +7

    cs.CVarXiv:2609.04196v12026
  60. A Simple Exponential Family Framework for Zero-Shot Learning

    Vinay Kumar Verma, Piyush Rai

    cs.LGcs.CVstat.MLarXiv:1707.08040v32017