Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,441 to 1,500 of 18,802

  1. Artificial Intelligence Literacy and Sustainable Development: An Ethical Governance and Development Goals Framework

    Md. Masudul Islam, Mirza Niaz Morshed, Md. Shafiqul Islam

    cs.CVarXiv:2609.10489v12026
  2. A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

    Changhao Chen, Bing Wang, Chris Xiaoxuan Lu +2

    cs.CVcs.LGcs.ROarXiv:2006.12567v22020
  3. Precision in Rice Variety Classification using Stacking-Based Ensemble Learning

    Md. Masudul Islam, Galib Muhammad Shahriar Himel, Md. Golam Moazzam +1

    cs.CVarXiv:2609.10524v12026
  4. Similar Image Search for Histopathology: SMILY

    Narayan Hegde, Jason D. Hipp, Yun Liu +11

    cs.CVq-bio.QMarXiv:1901.11112v32019
  5. MgSvF: Multi-Grained Slow vs. Fast Framework for Few-Shot Class-Incremental Learning

    Hanbin Zhao, Yongjian Fu, Mintong Kang +3

    cs.CVcs.LGarXiv:2006.15524v42020
  6. Region-aware Adaptive Instance Normalization for Image Harmonization

    Jun Ling, Han Xue, Li Song +2

    cs.CVarXiv:2106.02853v12021
  7. Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation

    Zhixiang Wei, Lin Chen, Yi Jin +6

    cs.CVarXiv:2312.04265v52023
  8. SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

    Athanasios Tragakis, Marco Aversa, Daniela Ivanova +4

    cs.CVcs.GRarXiv:2609.10363v12026
  9. Gen-LaneNet: A Generalized and Scalable Approach for 3D Lane Detection

    Yuliang Guo, Guang Chen, Peitao Zhao +4

    cs.CVarXiv:2003.10656v12020
  10. BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation Models

    Junfeng Xia, Wenhao Ye, Junxiang Zhang +3

    cs.CVq-bio.NCarXiv:2609.10518v12026
  11. Enhanced Deformable Convolution with Center-invariant Offset and Edge-aware Mask

    Yixiao Li, Xiaoyuan Yang, Jin Jiang +5

    cs.CVarXiv:2609.10387v12026
  12. Ensemble of Deep Convolutional Neural Networks for Automatic Pavement Crack Detection and Measurement

    Zhun Fan, Chong Li, Ying Chen +4

    cs.CVcs.LGeess.IVarXiv:2002.03241v12020
  13. AgroVisNet: A lightweight Convolutional Network and the BD-PlantDX Expert-Validated Benchmark for Radish, Potato and Pointed Gourd Disease Classification

    Md. Abdullah Mandal, Saad Ahmed, Md. Khalid Syfullah

    cs.CVarXiv:2609.10469v12026
  14. Advanced Brain Tissue Imaging with Data-Consistent Diffusion Priors in Laminographic X-Ray Nanoimaging

    Wenxuan Fang, Abraham L. Levitan, Ana Diaz +10

    cs.CVarXiv:2609.10456v12026
  15. Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation

    Suraj Nair, Chelsea Finn

    cs.LGcs.AIcs.CVarXiv:1909.05829v12019
  16. Dynamic Graph Message Passing Networks

    Li Zhang, Dan Xu, Anurag Arnab +1

    cs.CVcs.LGarXiv:1908.06955v52019
  17. Shape-guided Gaussian Splatting for Sparse-View X-ray 3D Reconstruction

    Pranav Poudel, Florence Dell'Aniello Picard, Nairouz Shehata +2

    cs.CVarXiv:2609.10376v12026
  18. Beyond Weak Labels: Prompt-Guided Local Refinement for Weakly Supervised Water Segmentation in High-Resolution Multispectral Imagery

    Muhammad Farhan Humayun, Mohammad Imangholiloo, Afifah Shah +2

    cs.CVarXiv:2609.10371v12026
  19. Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs

    Xuan Cuong Ngo, Ngan Le

    cs.CVarXiv:2609.10333v12026
  20. Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes

    Benedetta Liberatori, Nermin Samet, Paolo Rota +4

    cs.CVarXiv:2609.10356v12026
  21. Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells

    Vladimir Nekrasov, Hao Chen, Chunhua Shen +1

    cs.CVarXiv:1810.10804v32018
  22. IMAGDressing-v1: Customizable Virtual Dressing

    Fei Shen, Xin Jiang, Xin He +5

    cs.CVarXiv:2407.12705v22024
  23. Using U-Net Network for Efficient Brain Tumor Segmentation in MRI Images

    Jason Walsh, Alice Othmani, Mayank Jain +1

    eess.IVcs.CVq-bio.QMarXiv:2211.01885v12022
  24. Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval

    Xun Yang, Jianfeng Dong, Yixin Cao +3

    cs.CVarXiv:2007.02503v12020
  25. Improve Vision Language Model Chain-of-thought Reasoning

    Ruohong Zhang, Bowen Zhang, Yanghao Li +6

    cs.AIcs.CVarXiv:2410.16198v12024
  26. Adversarial Synthesis Learning Enables Segmentation Without Target Modality Ground Truth

    Yuankai Huo, Zhoubing Xu, Shunxing Bao +3

    cs.CVarXiv:1712.07695v12017
  27. Dimensionality Reduction for Hyperspectral Image Classification

    Mohamed Cherifi, Ammar Mesloub, Mohammed Nabil El Korso +2

    cs.CVeess.SParXiv:2609.10334v12026
  28. Graph-FCN for image semantic segmentation

    Yi Lu, Yaran Chen, Dongbin Zhao +1

    cs.CVarXiv:2001.00335v12020
  29. Kaolin: A PyTorch Library for Accelerating 3D Deep Learning Research

    Krishna Murthy Jatavallabhula, Edward Smith, Jean-Francois Lafleche +6

    cs.CVcs.LGcs.ROarXiv:1911.05063v22019
  30. Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition

    Diogo C Luvizon, Hedi Tabia, David Picard

    cs.CVarXiv:1912.08077v22019
  31. Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge

    Samed Doğan, Nico Leuze, Alfred Schöttl

    cs.CVarXiv:2609.10322v12026
  32. Decoupled Self-Forcing Distillation for Streaming Talking Head Generation

    Yanru An, Ruiyan Wang, Wenwu Wei +7

    cs.CVarXiv:2609.10317v12026
  33. IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks

    Youngwoon Lee, Edward S. Hu, Zhengyu Yang +2

    cs.ROcs.AIcs.CVarXiv:1911.07246v12019
  34. SynThermFace: Amplifying Limited Paired Data for Visible-Thermal Face Recognition via Synthetic Data Generation

    Anjith George, Adam Unal, Sebastien Marcel

    cs.CVarXiv:2609.10303v12026
  35. Deep Learning in Breast Cancer Imaging: A Decade of Progress and Future Directions

    Luyang Luo, Xi Wang, Yi Lin +7

    eess.IVcs.CVarXiv:2304.06662v42023
  36. Attention Driven Person Re-identification

    Fan Yang, Ke Yan, Shijian Lu +3

    cs.CVarXiv:1810.05866v12018
  37. Isotropic Embedding Perturbations for Robust Vision Language Encoders

    Hyesong Choi, Daeun Kim, Song Park +5

    cs.CVarXiv:2609.10292v12026
  38. ResNet or DenseNet? Introducing Dense Shortcuts to ResNet

    Chaoning Zhang, Philipp Benz, Dawit Mureja Argaw +5

    cs.CVarXiv:2010.12496v12020
  39. FreqFLD: Towards All-in-One Facial Landmark Detection via Frequency Modulation

    Shun Ren, Kaijie Jin, Shengkai Hu +5

    cs.CVarXiv:2609.10278v12026
  40. When Fusion Fails: Corruption-Aware Rebalanced Fusion for Multi-Modal Medical Image Segmentation

    Yuchen Pei, Xiaoyu Hu, Yixiong Zou +5

    cs.CVarXiv:2609.10261v12026
  41. Grid-guided Neural Radiance Fields for Large Urban Scenes

    Linning Xu, Yuanbo Xiangli, Sida Peng +5

    cs.CVarXiv:2303.14001v12023
  42. LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting

    Donghun Ryu, Minhyeok Lee

    cs.CVarXiv:2609.10095v12026
  43. Interpreting Object-Dependent Concept Brittleness in Text-to-Image Diffusion Models

    Yifan Yuan, Xiangyu Liu, Hongming Shan +5

    cs.CVcs.MMarXiv:2609.09909v12026
  44. Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

    Yi Tay, Mostafa Dehghani, Jinfeng Rao +7

    cs.CLcs.AIcs.CVarXiv:2109.10686v22021
  45. Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest

    Jack Hessel, Ana Marasović, Jena D. Hwang +5

    cs.CLcs.CVarXiv:2209.06293v22022
  46. DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting

    Agelos Kratimenos, Jiahui Lei, Kostas Daniilidis

    cs.CVcs.GRarXiv:2312.00112v22023
  47. Robust LSTM-Autoencoders for Face De-Occlusion in the Wild

    Fang Zhao, Jiashi Feng, Jian Zhao +2

    cs.CVarXiv:1612.08534v12016
  48. UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport

    Zonglin Yang, Huilan Ma, Xudan Zheng +1

    cs.CVarXiv:2609.10224v12026
  49. Knowledge Distillation via the Target-aware Transformer

    Sihao Lin, Hongwei Xie, Bing Wang +4

    cs.CVarXiv:2205.10793v22022
  50. Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields

    Jingbo Zhang, Xiaoyu Li, Ziyu Wan +2

    cs.CVcs.GRarXiv:2305.11588v22023
  51. 3rd Place Solution to Human Motion Challenges in Real-World and Clinical Settings (MoCha) @ECCV2026: Language-Aligned Motion Representations for Domain-Generalizable UPDRS-Gait Severity Estimation

    Soojie Kim, Muhammad Munsif, Minkyung Kim +1

    cs.CVarXiv:2609.10187v12026
  52. Deep Multitask Architecture for Integrated 2D and 3D Human Sensing

    Alin-Ionut Popa, Mihai Zanfir, Cristian Sminchisescu

    cs.CVarXiv:1701.08985v12017
  53. ScopeMamba-YOLO: Widening the Perceptual Scope Inward and Outward for Small Object Detection in Remote Sensing Imagery

    Junjie Fan, Yijun Mai, Linduo Wei +5

    cs.CVarXiv:2609.10156v12026
  54. CubeNet: Equivariance to 3D Rotation and Translation

    Daniel Worrall, Gabriel Brostow

    cs.CVcs.AIcs.LGarXiv:1804.04458v12018
  55. Dynamic Feature Integration for Simultaneous Detection of Salient Object, Edge and Skeleton

    Jiang-Jiang Liu, Qibin Hou, Ming-Ming Cheng

    cs.CVarXiv:2004.08595v12020
  56. TransGaze-Object: Transformer Based Driver Gaze Object Prediction Framework in Real Driving

    Pavan Kumar Sharma, Ayush Pande, Pranamesh Chakraborty

    cs.CVarXiv:2609.10139v12026
  57. Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

    Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar +2

    cs.CVarXiv:2609.10008v12026
  58. Feature-map-level Online Adversarial Knowledge Distillation

    Inseop Chung, SeongUk Park, Jangho Kim +1

    cs.LGcs.AIcs.CVarXiv:2002.01775v32020
  59. From Few-Shot Segmentation to Clinician-in-the-Loop Medical Image Analysis

    Yazhou Zhu

    cs.CVarXiv:2609.10001v12026
  60. Zero-Shot Recognition using Dual Visual-Semantic Mapping Paths

    Yanan Li, Donghui Wang, Huanhang Hu +2

    cs.CVarXiv:1703.05002v22017