Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,621 to 4,680 of 18,815

  1. TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening

    Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall +2

    cs.CVcs.LGarXiv:2609.00300v12026
  2. Unsupervised Cross-dataset Person Re-identification by Transfer Learning of Spatial-Temporal Patterns

    Jianming Lv, Weihang Chen, Qing Li +1

    cs.CVarXiv:1803.07293v12018
  3. Less Is More: Balancing Positive and Negative Space in Visual Concept Blending

    Shishi Xiao, Adam J. Coscia, David H. Laidlaw

    cs.CVcs.HCarXiv:2609.00476v12026
  4. Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

    Fizza Rubab, Yiying Tong, Arun Ross

    cs.CVarXiv:2609.00411v12026
  5. MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation

    Yuxiang Fu, Qi Yan, Lele Wang +2

    cs.CVcs.AIcs.LGarXiv:2503.09950v12025
  6. Weakly Supervised 3D Hand Pose Estimation via Biomechanical Constraints

    Adrian Spurr, Umar Iqbal, Pavlo Molchanov +2

    cs.CVarXiv:2003.09282v22020
  7. SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling

    Chad Wong, Sicheng Chen, Tianyi Zhang +4

    cs.CVarXiv:2609.00396v12026
  8. K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs

    Ziheng Ouyang, Zhen Li, Qibin Hou

    cs.CVarXiv:2502.18461v22025
  9. Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies

    Shaibal Saha, Lanyu Xu

    cs.CVcs.ARarXiv:2503.02891v32025
  10. Difference of Normals as a Multi-Scale Operator in Unorganized Point Clouds

    Yani Ioannou, Babak Taati, Robin Harrap +1

    cs.CVarXiv:1209.1759v12012
  11. SeedEdit 3.0: Fast and High-Quality Generative Image Editing

    Peng Wang, Yichun Shi, Xiaochen Lian +5

    cs.CVarXiv:2506.05083v22025
  12. Dataset Distillation via Factorization

    Songhua Liu, Kai Wang, Xingyi Yang +2

    cs.CVcs.LGarXiv:2210.16774v12022
  13. Membership is Ownership: A Robust Ownership Verification Framework for Diffusion Models

    Feng Jiang, Zuobin Xiong, An Huang +2

    cs.CRcs.CVarXiv:2608.28929v12026
  14. Online Coreset Selection for Rehearsal-based Continual Learning

    Jaehong Yoon, Divyam Madaan, Eunho Yang +1

    cs.LGcs.CVarXiv:2106.01085v42021
  15. Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

    Sunwhi Kim, Sunyul Kim, Meounggun Jo +1

    cs.HCcs.CVarXiv:2608.30210v12026
  16. GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

    Omid Taheri, Vasileios Choutas, Michael J. Black +1

    cs.CVarXiv:2112.11454v22021
  17. Domain Impression: A Source Data Free Domain Adaptation Method

    Vinod K Kurmi, Venkatesh K Subramanian, Vinay P Namboodiri

    cs.CVcs.AIarXiv:2102.09003v12021
  18. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

    Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian +4

    cs.ROcs.AIcs.CVarXiv:2601.03782v12026
  19. Deep Nearest Neighbor Anomaly Detection

    Liron Bergman, Niv Cohen, Yedid Hoshen

    cs.LGcs.CVstat.MLarXiv:2002.10445v12020
  20. ConCA: Concentration-Aware Channel Attention for Fine-Grained Visual Recognition

    Yu-Sheng Liu, Yu-Chen Tung

    hep-excs.CVarXiv:2608.30183v12026
  21. Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Video

    Hou-Ning Hu, Yen-Chen Lin, Ming-Yu Liu +3

    cs.CVcs.GRcs.MMarXiv:1705.01759v12017
  22. YOLOv12: A Breakdown of the Key Architectural Features

    Mujadded Al Rabbani Alif, Muhammad Hussain

    cs.CVcs.AIarXiv:2502.14740v12025
  23. Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

    Huilin Deng, Ding Zou, Rui Ma +3

    cs.CVarXiv:2503.07065v12025
  24. Action-Agnostic Human Pose Forecasting

    Hsu-kuang Chiu, Ehsan Adeli, Borui Wang +2

    cs.CVarXiv:1810.09676v12018
  25. Dual Mixup Regularized Learning for Adversarial Domain Adaptation

    Yuan Wu, Diana Inkpen, Ahmed El-Roby

    cs.LGcs.CVstat.MLarXiv:2007.03141v22020
  26. Explaining in Style: Training a GAN to explain a classifier in StyleSpace

    Oran Lang, Yossi Gandelsman, Michal Yarom +8

    cs.CVcs.LGcs.NEarXiv:2104.13369v22021
  27. CyCLIP: Cyclic Contrastive Language-Image Pretraining

    Shashank Goel, Hritik Bansal, Sumit Bhatia +3

    cs.CVcs.LGarXiv:2205.14459v22022
  28. DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation

    Bo-Wen Yin, Jiao-Long Cao, Ming-Ming Cheng +1

    cs.CVarXiv:2504.04701v12025
  29. VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing

    Xiangpeng Yang, Linchao Zhu, Hehe Fan +1

    cs.CVarXiv:2502.17258v12025
  30. Towards a Unified Copernicus Foundation Model for Earth Vision

    Yi Wang, Zhitong Xiong, Chenying Liu +8

    cs.CVarXiv:2503.11849v32025
  31. GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

    Zhiyuan Yan, Junyan Ye, Weijia Li +7

    cs.CVarXiv:2504.02782v32025
  32. Seeing the Unseen: Camouflaged Object Detection Beyond the Visible Spectrum

    Avi Gupta, Trasha Gupta

    cs.CVarXiv:2608.30355v12026
  33. From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer

    Haipeng Xiong, Hao Lu, Chengxin Liu +3

    cs.CVarXiv:1908.06473v12019
  34. ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

    Jiawei Zhang, Hongsong Wang, Pan Zhou

    cs.CVcs.AIarXiv:2608.30307v12026
  35. Reconstruction-Aware Cryo-EM Particle Picking

    Riku Itsuji, Yuanhao Wang, Xingjian Li +3

    q-bio.BMcs.CVarXiv:2608.28838v12026
  36. Generative Translation Priors: Bayesian Imaging with Cross-Modality Image Translation

    Evan Bell, Jiaming Liu, Yifan Chen +1

    eess.IVcs.CVcs.LGarXiv:2608.28872v12026
  37. Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation

    Kang Liu, Zhuoqi Ma, Xiaolu Kang +4

    cs.CVcs.AIarXiv:2502.20056v12025
  38. Learning Dynamic Routing for Semantic Segmentation

    Yanwei Li, Lin Song, Yukang Chen +4

    cs.CVarXiv:2003.10401v12020
  39. Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer

    Wang Zeng, Sheng Jin, Wentao Liu +4

    cs.CVarXiv:2204.08680v32022
  40. UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum and Iterative Generalist-Specialist Learning

    Weikang Wan, Haoran Geng, Yun Liu +4

    cs.ROcs.CVarXiv:2304.00464v22023
  41. Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression

    Yeying Jin, Wenhan Yang, Robby T. Tan

    cs.CVarXiv:2207.10564v22022
  42. Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation

    Kailin Li, Zhenxin Li, Shiyi Lan +6

    cs.CVarXiv:2503.12820v12025
  43. CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection

    Youwei Pang, Xiaoqi Zhao, Lihe Zhang +1

    cs.CVarXiv:2112.02363v32021
  44. Medical Foundation Model Features as Perceptual Loss for Brain MRI Contrast Dose Simulation

    Changsheng Fang, Dayang Wang, T. Campbell Arnold +2

    eess.IVcs.CVarXiv:2608.28773v12026
  45. Incremental Few-Shot Learning with Attention Attractor Networks

    Mengye Ren, Renjie Liao, Ethan Fetaya +1

    cs.LGcs.CVstat.MLarXiv:1810.07218v32018
  46. ADN: Artifact Disentanglement Network for Unsupervised Metal Artifact Reduction

    Haofu Liao, Wei-An Lin, S. Kevin Zhou +1

    eess.IVcs.CVarXiv:1908.01104v42019
  47. Domain Adaptation with Auxiliary Target Domain-Oriented Classifier

    Jian Liang, Dapeng Hu, Jiashi Feng

    cs.CVcs.LGarXiv:2007.04171v52020
  48. HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos

    Jinglei Zhang, Jiankang Deng, Chao Ma +1

    cs.CVarXiv:2501.02973v12025
  49. Building and better understanding vision-language models: insights and future directions

    Hugo Laurençon, Andrés Marafioti, Victor Sanh +1

    cs.CVcs.AIarXiv:2408.12637v12024
  50. Multi-subject Open-set Personalization in Video Generation

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace +7

    cs.CVarXiv:2501.06187v22025
  51. MarioNETte: Few-shot Face Reenactment Preserving Identity of Unseen Targets

    Sungjoo Ha, Martin Kersner, Beomsu Kim +2

    cs.CVarXiv:1911.08139v12019
  52. Facial expression and attributes recognition based on multi-task learning of lightweight neural networks

    Andrey V. Savchenko

    cs.CVarXiv:2103.17107v32021
  53. Scaling Language-Free Visual Representation Learning

    David Fan, Shengbang Tong, Jiachen Zhu +8

    cs.CVarXiv:2504.01017v12025
  54. Deep Learning for Screening COVID-19 using Chest X-Ray Images

    Sanhita Basu, Sushmita Mitra, Nilanjan Saha

    eess.IVcs.CVcs.LGarXiv:2004.10507v42020
  55. Back to the Features: DINO as a Foundation for Video World Models

    Federico Baldassarre, Marc Szafraniec, Basile Terver +6

    cs.CVarXiv:2507.19468v12025
  56. Improving neural implicit surfaces geometry with patch warping

    François Darmon, Bénédicte Bascle, Jean-Clément Devaux +2

    cs.CVarXiv:2112.09648v22021
  57. Flexible Isosurface Extraction for Gradient-Based Mesh Optimization

    Tianchang Shen, Jacob Munkberg, Jon Hasselgren +7

    cs.GRcs.CVcs.LGarXiv:2308.05371v12023
  58. Wavelength-based Attributed Deep Neural Network for Underwater Image Restoration

    Prasen Kumar Sharma, Ira Bisht, Arijit Sur

    eess.IVcs.CVarXiv:2106.07910v32021
  59. Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

    Guo Chen, Zhiqi Li, Shihao Wang +18

    cs.CVarXiv:2504.15271v12025
  60. Focal Inverse Distance Transform Maps for Crowd Localization

    Dingkang Liang, Wei Xu, Yingying Zhu +1

    cs.CVarXiv:2102.07925v32021