Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
7,141 to 7,200 of 18,817
What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models
Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3
cs.CVarXiv:2609.01551v12026Extreme View Synthesis
Inchang Choi, Orazio Gallo, Alejandro Troccoli +2
cs.CVarXiv:1812.04777v22018PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
Xiaolong Ma, Fu-Ming Guo, Wei Niu +5
cs.LGcs.CVcs.DCarXiv:1909.05073v42019Editable Visual Design
Junyan Ye, Wei Liu, Dongzhi Jiang +9
cs.CVcs.CLarXiv:2609.04034v12026Introduction to the Bag of Features Paradigm for Image Classification and Retrieval
Stephen O'Hara, Bruce A. Draper
cs.CVcs.IRarXiv:1101.3354v12011Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
Xiaojiang Peng, Limin Wang, Xingxing Wang +1
cs.CVarXiv:1405.4506v12014Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness
Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta +2
cs.CVcs.HCarXiv:2609.02512v12026AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates
Ning Liu, Xiaolong Ma, Zhiyuan Xu +3
cs.LGcs.AIcs.CVarXiv:1907.03141v22019Learning to Generate Images with Perceptual Similarity Metrics
Jake Snell, Karl Ridgeway, Renjie Liao +3
cs.LGcs.CVarXiv:1511.06409v32015Structured Prediction Helps 3D Human Motion Modelling
Emre Aksan, Manuel Kaufmann, Otmar Hilliges
cs.CVarXiv:1910.09070v12019CenterFormer: Center-based Transformer for 3D Object Detection
Zixiang Zhou, Xiangchen Zhao, Yu Wang +2
cs.CVarXiv:2209.05588v12022Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
Youmin Zhang, Yimin Chen, Xiao Bai +4
cs.CVarXiv:1909.03751v22019Evaluating the Impact of Intensity Normalization on MR Image Synthesis
Jacob C. Reinhold, Blake E. Dewey, Aaron Carass +1
cs.CVarXiv:1812.04652v12018Real-time Driver Drowsiness Detection for Android Application Using Deep Neural Networks Techniques
Rateb Jabbar, Khalifa Al-Khalifa, Mohamed Kharbeche +3
cs.CVcs.HCarXiv:1811.01627v12018Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding
Shanji Liu, Kelu Yao, Junxiao Xue +5
cs.CVarXiv:2609.02187v12026Human-centric Indoor Scene Synthesis Using Stochastic Grammar
Siyuan Qi, Yixin Zhu, Siyuan Huang +2
cs.CVarXiv:1808.08473v12018KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection
Zhouyuan Huai, Meiqi Wan, Yan Yang +4
cs.CVarXiv:2609.02077v12026Looking Beyond Appearances: Synthetic Training Data for Deep CNNs in Re-identification
Igor Barros Barbosa, Marco Cristani, Barbara Caputo +2
cs.CVarXiv:1701.03153v22017CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation
Hao Sun, Hongyi Wang, Jiaqing Liu +2
cs.MMcs.CLcs.CVarXiv:2207.14087v32022Low Frequency Adversarial Perturbation
Chuan Guo, Jared S. Frank, Kilian Q. Weinberger
cs.CVarXiv:1809.08758v22018Learning Semantic-Aware Knowledge Guidance for Low-Light Image Enhancement
Yuhui Wu, Chen Pan, Guoqing Wang +4
cs.CVarXiv:2304.07039v12023A Neural Temporal Model for Human Motion Prediction
Anand Gopalakrishnan, Ankur Mali, Dan Kifer +2
cs.CVarXiv:1809.03036v52018ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
Jitai Hao, Ke Yang, Qiang Huang +1
cs.CVcs.CLarXiv:2609.02780v12026Cross-Dataset Person Re-Identification via Unsupervised Pose Disentanglement and Adaptation
Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin +1
cs.CVarXiv:1909.09675v12019TACO: Trash Annotations in Context for Litter Detection
Pedro F Proença, Pedro Simões
cs.CVarXiv:2003.06975v22020CLIP-Driven Fine-grained Text-Image Person Re-identification
Shuanglin Yan, Neng Dong, Liyan Zhang +1
cs.CVarXiv:2210.10276v12022SpatialBot: Precise Spatial Understanding with Vision Language Models
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4
cs.CVarXiv:2406.13642v72024ELEGANT: Exchanging Latent Encodings with GAN for Transferring Multiple Face Attributes
Taihong Xiao, Jiapeng Hong, Jinwen Ma
cs.CVarXiv:1803.10562v22018Pose Guided Structured Region Ensemble Network for Cascaded Hand Pose Estimation
Xinghao Chen, Guijin Wang, Hengkai Guo +1
cs.CVarXiv:1708.03416v22017Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Keen You, Haotian Zhang, Eldon Schoop +5
cs.CVcs.CLcs.HCarXiv:2404.05719v12024SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation
Qin Lu, Zheyang Jing, Yujie Yang +3
cs.CVcs.AIarXiv:2609.02247v12026CASIA-SURF: A Large-scale Multi-modal Benchmark for Face Anti-spoofing
Shifeng Zhang, Ajian Liu, Jun Wan +5
cs.CVarXiv:1908.10654v22019Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation
Xiaoxiao Li, Chen Change Loy
cs.CVarXiv:1803.04242v22018Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models
Haoning Wu, Zicheng Zhang, Erli Zhang +11
cs.CVcs.MMarXiv:2311.06783v12023HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Zhaorun Chen, Zhuokai Zhao, Hongyin Luo +3
cs.CVcs.AIcs.LGarXiv:2403.00425v22024Map-Guided Curriculum Domain Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation
Christos Sakaridis, Dengxin Dai, Luc Van Gool
cs.CVarXiv:2005.14553v22020Summaries:한국어VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM
Sangmin Song, Sarath Kodagoda, Marc G. Carmichael +4
cs.CVcs.AIarXiv:2609.00775v12026FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age
Kimmo Kärkkäinen, Jungseock Joo
cs.CVcs.LGarXiv:1908.04913v12019Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain
Daizong Liu, Junhao Dong, Zhiyuan Ma +6
cs.CVarXiv:2609.00788v12026Real-time Cardiovascular MR with Spatio-temporal Artifact Suppression using Deep Learning - Proof of Concept in Congenital Heart Disease
Andreas Hauptmann, Simon Arridge, Felix Lucka +2
cs.CVcs.NEarXiv:1803.05192v32018Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification
Lin Wu, Yang Wang, Junbin Gao +1
cs.CVarXiv:1808.01911v22018RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing
Kaiyue Kang, Qixuan He, Peijin Wang +9
cs.CVarXiv:2609.00814v12026Abnormality Detection and Localization in Chest X-Rays using Deep Convolutional Neural Networks
Mohammad Tariqul Islam, Md Abdul Aowal, Ahmed Tahseen Minhaz +1
cs.CVarXiv:1705.09850v32017Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison
Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit
cs.CVarXiv:2609.01530v12026HorizonNet: Learning Room Layout with 1D Representation and Pano Stretch Data Augmentation
Cheng Sun, Chi-Wei Hsiao, Min Sun +1
cs.CVarXiv:1901.03861v22019OVANet: One-vs-All Network for Universal Domain Adaptation
Kuniaki Saito, Kate Saenko
cs.CVarXiv:2104.03344v42021ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence
Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao +1
cs.CVarXiv:2609.01344v12026Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot Learning
Zhiqiang Shen, Zechun Liu, Jie Qin +2
cs.CVcs.AIcs.LGarXiv:2102.03983v12021Panoptic NeRF: 3D-to-2D Label Transfer for Panoptic Urban Scene Segmentation
Xiao Fu, Shangzhan Zhang, Tianrun Chen +5
cs.CVarXiv:2203.15224v22022TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views
Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9
cs.CVarXiv:2609.01899v12026Fourier Space Losses for Efficient Perceptual Image Super-Resolution
Dario Fuoli, Luc Van Gool, Radu Timofte
eess.IVcs.CVarXiv:2106.00783v12021Deep Learning for Human Affect Recognition: Insights and New Developments
Philipp V. Rouast, Marc T. P. Adam, Raymond Chiong
cs.LGcs.AIcs.CVarXiv:1901.02884v12019A Survey of Deep Learning for Mathematical Reasoning
Pan Lu, Liang Qiu, Wenhao Yu +2
cs.AIcs.CLcs.CVarXiv:2212.10535v22022Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment
Mohamad Jouni, Aurélien Godet, Mauro Dalla Mura
eess.IVcs.CVarXiv:2609.01060v12026Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study
Samuel Ritter, David G. T. Barrett, Adam Santoro +1
stat.MLcs.CVcs.LGarXiv:1706.08606v22017UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations
Chad Melton, PhD., Annabelle Kelton
cs.CVcs.DBarXiv:2609.01738v12026Loss Aware Post-training Quantization
Yury Nahshan, Brian Chmiel, Chaim Baskin +4
cs.LGcs.CVarXiv:1911.07190v22019One-pass Multi-task Networks with Cross-task Guided Attention for Brain Tumor Segmentation
Chenhong Zhou, Changxing Ding, Xinchao Wang +2
cs.CVcs.AIcs.LGarXiv:1906.01796v22019Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Kang Liao, Yihang Luo, Xiao-Ming Wu +7
cs.CVarXiv:2609.04196v12026A Simple Exponential Family Framework for Zero-Shot Learning
Vinay Kumar Verma, Piyush Rai
cs.LGcs.CVstat.MLarXiv:1707.08040v32017