Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,501 to 1,560 of 18,848
Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
Suraj Nair, Chelsea Finn
cs.LGcs.AIcs.CVarXiv:1909.05829v12019Dynamic Graph Message Passing Networks
Li Zhang, Dan Xu, Anurag Arnab +1
cs.CVcs.LGarXiv:1908.06955v52019Shape-guided Gaussian Splatting for Sparse-View X-ray 3D Reconstruction
Pranav Poudel, Florence Dell'Aniello Picard, Nairouz Shehata +2
cs.CVarXiv:2609.10376v12026Beyond Weak Labels: Prompt-Guided Local Refinement for Weakly Supervised Water Segmentation in High-Resolution Multispectral Imagery
Muhammad Farhan Humayun, Mohammad Imangholiloo, Afifah Shah +2
cs.CVarXiv:2609.10371v12026Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs
Xuan Cuong Ngo, Ngan Le
cs.CVarXiv:2609.10333v12026Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes
Benedetta Liberatori, Nermin Samet, Paolo Rota +4
cs.CVarXiv:2609.10356v12026Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells
Vladimir Nekrasov, Hao Chen, Chunhua Shen +1
cs.CVarXiv:1810.10804v32018IMAGDressing-v1: Customizable Virtual Dressing
Fei Shen, Xin Jiang, Xin He +5
cs.CVarXiv:2407.12705v22024Using U-Net Network for Efficient Brain Tumor Segmentation in MRI Images
Jason Walsh, Alice Othmani, Mayank Jain +1
eess.IVcs.CVq-bio.QMarXiv:2211.01885v12022Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval
Xun Yang, Jianfeng Dong, Yixin Cao +3
cs.CVarXiv:2007.02503v12020Improve Vision Language Model Chain-of-thought Reasoning
Ruohong Zhang, Bowen Zhang, Yanghao Li +6
cs.AIcs.CVarXiv:2410.16198v12024Adversarial Synthesis Learning Enables Segmentation Without Target Modality Ground Truth
Yuankai Huo, Zhoubing Xu, Shunxing Bao +3
cs.CVarXiv:1712.07695v12017Dimensionality Reduction for Hyperspectral Image Classification
Mohamed Cherifi, Ammar Mesloub, Mohammed Nabil El Korso +2
cs.CVeess.SParXiv:2609.10334v12026Graph-FCN for image semantic segmentation
Yi Lu, Yaran Chen, Dongbin Zhao +1
cs.CVarXiv:2001.00335v12020Kaolin: A PyTorch Library for Accelerating 3D Deep Learning Research
Krishna Murthy Jatavallabhula, Edward Smith, Jean-Francois Lafleche +6
cs.CVcs.LGcs.ROarXiv:1911.05063v22019Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
Diogo C Luvizon, Hedi Tabia, David Picard
cs.CVarXiv:1912.08077v22019Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge
Samed Doğan, Nico Leuze, Alfred Schöttl
cs.CVarXiv:2609.10322v12026Decoupled Self-Forcing Distillation for Streaming Talking Head Generation
Yanru An, Ruiyan Wang, Wenwu Wei +7
cs.CVarXiv:2609.10317v12026IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks
Youngwoon Lee, Edward S. Hu, Zhengyu Yang +2
cs.ROcs.AIcs.CVarXiv:1911.07246v12019SynThermFace: Amplifying Limited Paired Data for Visible-Thermal Face Recognition via Synthetic Data Generation
Anjith George, Adam Unal, Sebastien Marcel
cs.CVarXiv:2609.10303v12026Deep Learning in Breast Cancer Imaging: A Decade of Progress and Future Directions
Luyang Luo, Xi Wang, Yi Lin +7
eess.IVcs.CVarXiv:2304.06662v42023Attention Driven Person Re-identification
Fan Yang, Ke Yan, Shijian Lu +3
cs.CVarXiv:1810.05866v12018Isotropic Embedding Perturbations for Robust Vision Language Encoders
Hyesong Choi, Daeun Kim, Song Park +5
cs.CVarXiv:2609.10292v12026ResNet or DenseNet? Introducing Dense Shortcuts to ResNet
Chaoning Zhang, Philipp Benz, Dawit Mureja Argaw +5
cs.CVarXiv:2010.12496v12020FreqFLD: Towards All-in-One Facial Landmark Detection via Frequency Modulation
Shun Ren, Kaijie Jin, Shengkai Hu +5
cs.CVarXiv:2609.10278v12026When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation
Yuchen Pei, Xiaoyu Hu, Yixiong Zou +5
cs.CVarXiv:2609.10261v12026Grid-guided Neural Radiance Fields for Large Urban Scenes
Linning Xu, Yuanbo Xiangli, Sida Peng +5
cs.CVarXiv:2303.14001v12023LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting
Donghun Ryu, Minhyeok Lee
cs.CVarXiv:2609.10095v12026Interpreting Object-Dependent Concept Brittleness in Text-to-Image Diffusion Models
Yifan Yuan, Xiangyu Liu, Hongming Shan +5
cs.CVcs.MMarXiv:2609.09909v12026Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
Yi Tay, Mostafa Dehghani, Jinfeng Rao +7
cs.CLcs.AIcs.CVarXiv:2109.10686v22021Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
Jack Hessel, Ana Marasović, Jena D. Hwang +5
cs.CLcs.CVarXiv:2209.06293v22022DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting
Agelos Kratimenos, Jiahui Lei, Kostas Daniilidis
cs.CVcs.GRarXiv:2312.00112v22023Robust LSTM-Autoencoders for Face De-Occlusion in the Wild
Fang Zhao, Jiashi Feng, Jian Zhao +2
cs.CVarXiv:1612.08534v12016UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport
Zonglin Yang, Huilan Ma, Xudan Zheng +1
cs.CVarXiv:2609.10224v12026Knowledge Distillation via the Target-aware Transformer
Sihao Lin, Hongwei Xie, Bing Wang +4
cs.CVarXiv:2205.10793v22022Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
Jingbo Zhang, Xiaoyu Li, Ziyu Wan +2
cs.CVcs.GRarXiv:2305.11588v220233rd Place Solution to Human Motion Challenges in Real-World and Clinical Settings (MoCha) @ECCV2026: Language-Aligned Motion Representations for Domain-Generalizable UPDRS-Gait Severity Estimation
Soojie Kim, Muhammad Munsif, Minkyung Kim +1
cs.CVarXiv:2609.10187v12026Deep Multitask Architecture for Integrated 2D and 3D Human Sensing
Alin-Ionut Popa, Mihai Zanfir, Cristian Sminchisescu
cs.CVarXiv:1701.08985v12017ScopeMamba-YOLO: Widening the Perceptual Scope Inward and Outward for Small Object Detection in Remote Sensing Imagery
Junjie Fan, Yijun Mai, Linduo Wei +5
cs.CVarXiv:2609.10156v12026CubeNet: Equivariance to 3D Rotation and Translation
Daniel Worrall, Gabriel Brostow
cs.CVcs.AIcs.LGarXiv:1804.04458v12018Dynamic Feature Integration for Simultaneous Detection of Salient Object, Edge and Skeleton
Jiang-Jiang Liu, Qibin Hou, Ming-Ming Cheng
cs.CVarXiv:2004.08595v12020TransGaze-Object: Transformer Based Driver Gaze Object Prediction Framework in Real Driving
Pavan Kumar Sharma, Ayush Pande, Pranamesh Chakraborty
cs.CVarXiv:2609.10139v12026Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval
Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar +2
cs.CVarXiv:2609.10008v12026Feature-map-level Online Adversarial Knowledge Distillation
Inseop Chung, SeongUk Park, Jangho Kim +1
cs.LGcs.AIcs.CVarXiv:2002.01775v32020From Few-Shot Segmentation to Clinician-in-the-Loop Medical Image Analysis
Yazhou Zhu
cs.CVarXiv:2609.10001v12026Zero-Shot Recognition using Dual Visual-Semantic Mapping Paths
Yanan Li, Donghui Wang, Huanhang Hu +2
cs.CVarXiv:1703.05002v22017Fusion of convolution neural network, support vector machine and Sobel filter for accurate detection of COVID-19 patients using X-ray images
Danial Sharifrazi, Roohallah Alizadehsani, Mohamad Roshanzamir +13
eess.IVcs.CVarXiv:2102.06883v12021Learning to Predict Vehicle Trajectories with Model-based Planning
Haoran Song, Di Luan, Wenchao Ding +2
cs.CVcs.ROarXiv:2103.04027v22021Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
Zizhen Wang, Bo Feng, Zhengfeng Lai +5
cs.CVarXiv:2609.09973v12026Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry
Oriol Marín, Roger Marí, Gloria Haro +1
cs.CVarXiv:2609.09924v12026Efficient and Degradation-Adaptive Network for Real-World Image Super-Resolution
Jie Liang, Hui Zeng, Lei Zhang
cs.CVeess.IVarXiv:2203.14216v12022Reverse Classification Accuracy: Predicting Segmentation Performance in the Absence of Ground Truth
Vanya V. Valindria, Ioannis Lavdas, Wenjia Bai +5
cs.CVarXiv:1702.03407v12017Efficient Visual Tracking with Exemplar Transformers
Philippe Blatter, Menelaos Kanakis, Martin Danelljan +1
cs.CVarXiv:2112.09686v42021ST-LLM: Large Language Models Are Effective Temporal Learners
Ruyang Liu, Chen Li, Haoran Tang +3
cs.CVarXiv:2404.00308v12024CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
Christian Diller, Angela Dai
cs.CVarXiv:2311.16097v22023From Pixels to Hierarchical Sequences: Quadtree Mask Encoding for Vision-Language Binary Change Detection
Xiao An, Ruikang Zhang, Chen Zhong +4
cs.CVarXiv:2609.09876v12026Temporal Action Localization by Structured Maximal Sums
Zehuan Yuan, Jonathan C. Stroud, Tong Lu +1
cs.CVarXiv:1704.04671v12017Pretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification
Philip Graemer, Giuseppe Di Caprio
cs.CVarXiv:2609.09863v12026Category Level Object Pose Estimation via Neural Analysis-by-Synthesis
Xu Chen, Zijian Dong, Jie Song +2
cs.CVarXiv:2008.08145v12020SkNeXt enables topology-guided neuronal reconstruction from petabyte-scale microscopy data
Jiayi Ding, Hu Zhao
cs.CVarXiv:2609.09832v12026