Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,641 to 5,700 of 18,855

  1. Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set

    Michael Zang, Haiyu Wu, Mrinal Sharma +1

    cs.CVcs.AIarXiv:2609.01141v12026
  2. Modulated Periodic Activations for Generalizable Local Functional Representations

    Ishit Mehta, Michaël Gharbi, Connelly Barnes +3

    cs.CVcs.GRarXiv:2104.03960v12021
    Summaries:한국어
  3. LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

    Chenming Zhu, Tai Wang, Wenwei Zhang +2

    cs.CVarXiv:2409.18125v32024
  4. Cameras as Relative Positional Encoding

    Ruilong Li, Brent Yi, Junchen Liu +3

    cs.CVcs.AIarXiv:2507.10496v22025
  5. VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

    Qi Wang, Yanrui Yu, Ye Yuan +2

    cs.CVarXiv:2505.12434v42025
  6. Localizing Anomalies from Weakly-Labeled Videos

    Hui Lv, Chuanwei Zhou, Chunyan Xu +2

    cs.CVarXiv:2008.08944v32020
  7. Challenge-Aware RGBT Tracking

    Chenglong Li, Lei Liu, Andong Lu +2

    cs.CVarXiv:2007.13143v12020
  8. Self-Rewarding Vision-Language Model via Reasoning Decomposition

    Zongxia Li, Wenhao Yu, Chengsong Huang +8

    cs.CVarXiv:2508.19652v22025
  9. Generalization in diffusion models arises from geometry-adaptive harmonic representations

    Zahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli +1

    cs.CVcs.LGarXiv:2310.02557v32023
  10. Probing the 3D Awareness of Visual Foundation Models

    Mohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis +7

    cs.CVarXiv:2404.08636v12024
  11. Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

    Zhixuan Liang, Yizhuo Li, Tianshuo Yang +9

    cs.CVcs.LGcs.ROarXiv:2508.20072v42025
  12. VITA: Towards Open-Source Interactive Omni Multimodal LLM

    Chaoyou Fu, Haojia Lin, Zuwei Long +16

    cs.CVcs.AIcs.CLarXiv:2408.05211v32024
  13. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    HunyuanWorld Team, Zhenwei Wang, Yuhao Liu +52

    cs.CVarXiv:2507.21809v22025
  14. Learning-based Video Motion Magnification

    Tae-Hyun Oh, Ronnachai Jaroensri, Changil Kim +4

    cs.CVcs.GRarXiv:1804.02684v32018
  15. InstantAvatar: Learning Avatars from Monocular Video in 60 Seconds

    Tianjian Jiang, Xu Chen, Jie Song +1

    cs.CVarXiv:2212.10550v12022
  16. Interactive Post-Training for Vision-Language-Action Models

    Shuhan Tan, Kairan Dou, Yue Zhao +1

    cs.LGcs.AIcs.CVarXiv:2505.17016v12025
  17. Self-supervised Video Object Segmentation by Motion Grouping

    Charig Yang, Hala Lamdouar, Erika Lu +2

    cs.CVcs.LGarXiv:2104.07658v22021
  18. Denoising Diffusion Bridge Models

    Linqi Zhou, Aaron Lou, Samar Khanna +1

    cs.CVcs.AIarXiv:2309.16948v32023
  19. UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

    Cong Wei, Yang Chen, Haonan Chen +5

    cs.CVcs.AIcs.CLarXiv:2311.17136v12023
  20. Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

    Tianyu Huang, Wangguandong Zheng, Tengfei Wang +8

    cs.CVarXiv:2506.04225v12025
  21. Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Yi Xin, Qi Qin, Siqi Luo +29

    cs.CVarXiv:2510.06308v12025
  22. TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection

    Kyle Min, Jason J. Corso

    cs.CVarXiv:1908.05786v12019
  23. RealCAD: Towards Real-World Image-to-CAD Reconstruction under Domain Shift and Parameter Bias

    Yihe Sun, Ziyu Lu, Kaihua Tang +1

    cs.CVarXiv:2608.30617v12026
  24. RELIC: Interactive Video World Model with Long-Horizon Memory

    Yicong Hong, Yiqun Mei, Chongjian Ge +11

    cs.CVarXiv:2512.04040v12025
  25. A Large Scale Event-based Detection Dataset for Automotive

    Pierre de Tournemire, Davide Nitti, Etienne Perot +2

    cs.CVcs.LGcs.ROarXiv:2001.08499v32020
  26. Latent Visual Reasoning

    Bangzheng Li, Ximeng Sun, Jiang Liu +7

    cs.CVcs.CLarXiv:2509.24251v22025
  27. DCFace: Synthetic Face Generation with Dual Condition Diffusion Model

    Minchul Kim, Feng Liu, Anil Jain +1

    cs.CVarXiv:2304.07060v12023
  28. Fast Video Generation with Sliding Tile Attention

    Peiyuan Zhang, Yongqi Chen, Runlong Su +4

    cs.CVarXiv:2502.04507v32025
  29. LayoutTransformer: Layout Generation and Completion with Self-attention

    Kamal Gupta, Justin Lazarow, Alessandro Achille +3

    cs.CVcs.LGarXiv:2006.14615v22020
  30. OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

    Gaojie Lin, Jianwen Jiang, Jiaqi Yang +2

    cs.CVarXiv:2502.01061v32025
  31. TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

    Xiaoxuan He, Siming Fu, Yuke Zhao +5

    cs.CVarXiv:2508.04324v42025
  32. Street Scene: A new dataset and evaluation protocol for video anomaly detection

    Bharathkumar Ramachandra, Michael Jones

    cs.CVarXiv:1902.05872v32019
  33. DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving

    Xiaosong Jia, Junqi You, Zhiyuan Zhang +1

    cs.LGcs.CVcs.ROarXiv:2503.07656v22025
  34. A Richly Annotated Dataset for Pedestrian Attribute Recognition

    Dangwei Li, Zhang Zhang, Xiaotang Chen +2

    cs.CVarXiv:1603.07054v32016
  35. GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

    Zebin Xing, Xingyu Zhang, Yang Hu +5

    cs.CVarXiv:2503.05689v62025
  36. Alleviating Over-segmentation Errors by Detecting Action Boundaries

    Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki +1

    cs.CVarXiv:2007.06866v12020
  37. Natural Image Matting via Guided Contextual Attention

    Yaoyi Li, Hongtao Lu

    cs.CVarXiv:2001.04069v12020
  38. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

    Shaoyuan Xie, Lingdong Kong, Yuhao Dong +5

    cs.CVcs.ROarXiv:2501.04003v12025
  39. Temporal-Relational CrossTransformers for Few-Shot Action Recognition

    Toby Perrett, Alessandro Masullo, Tilo Burghardt +2

    cs.CVarXiv:2101.06184v32021
  40. PixelDiT: Pixel Diffusion Transformers for Image Generation

    Yongsheng Yu, Wei Xiong, Weili Nie +3

    cs.CVarXiv:2511.20645v22025
  41. PARTFIELD: Learning 3D Feature Fields for Part Segmentation and Beyond

    Minghua Liu, Mikaela Angelina Uy, Donglai Xiang +4

    cs.CVarXiv:2504.11451v12025
  42. AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP

    Wenxin Ma, Xu Zhang, Qingsong Yao +6

    cs.CVcs.AIarXiv:2503.06661v12025
  43. Evidence-Guided Detection, Localization and Explanation for Text-Centric Image Forensics

    Peifeng Liu, Bin Li, Qingsong Zhang +3

    cs.CVarXiv:2609.02097v12026
  44. Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition

    Naoto Nishida, Yoshio Ishiguro

    cs.CVcs.HCcs.LGarXiv:2609.02510v12026
  45. HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

    Yi Li, Yuquan Deng, Jesse Zhang +9

    cs.ROcs.AIcs.CVarXiv:2502.05485v42025
  46. Learning Spatial-Aware Regressions for Visual Tracking

    Chong Sun, Dong Wang, Huchuan Lu +1

    cs.CVarXiv:1706.07457v22017
  47. Learning Deep Networks from Noisy Labels with Dropout Regularization

    Ishan Jindal, Matthew Nokleby, Xuewen Chen

    cs.CVcs.LGstat.MLarXiv:1705.03419v12017
  48. Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT

    Zhenyu Bu, Haoyan Ding, Chushu Shen +7

    cs.CVarXiv:2609.00447v12026
  49. Monet: Reasoning in Latent Visual Space Beyond Images and Language

    Qixun Wang, Yang Shi, Yifei Wang +5

    cs.CVcs.AIarXiv:2511.21395v22025
  50. Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

    Shuyao Xiao, Shengling Wang, Haoyu Niu +4

    cs.CVarXiv:2609.02000v12026
  51. A Survey on Deep Learning for Neuroimaging-based Brain Disorder Analysis

    Li Zhang, Mingliang Wang, Mingxia Liu +1

    eess.IVcs.CVcs.LGarXiv:2005.04573v12020
  52. Efficient Active Learning for Image Classification and Segmentation using a Sample Selection and Conditional Generative Adversarial Network

    Dwarikanath Mahapatra, Behzad Bozorgtabar, Jean-Philippe Thiran +1

    cs.CVarXiv:1806.05473v42018
  53. Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis

    Siddhant Shete, Hilmi Dogu Kücüker, Udo Frese +1

    cs.ROcs.ARcs.CVarXiv:2609.02219v12026
  54. Test-Time Training Done Right

    Tianyuan Zhang, Sai Bi, Yicong Hong +6

    cs.LGcs.CLcs.CVarXiv:2505.23884v12025
  55. Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image Classification

    Wen-Shuai Hu, Heng-Chao Li, Lei Pan +3

    cs.CVarXiv:1905.03577v22019
  56. Wan-Animate: Unified Character Animation and Replacement with Holistic Replication

    Gang Cheng, Xin Gao, Li Hu +23

    cs.CVarXiv:2509.14055v12025
  57. FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

    Roman Bachmann, Jesse Allardice, David Mizrahi +6

    cs.CVcs.LGarXiv:2502.13967v22025
  58. UniVideo: Unified Understanding, Generation, and Editing for Videos

    Cong Wei, Quande Liu, Zixuan Ye +5

    cs.CVarXiv:2510.08377v42025
  59. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    Tianwei Lin, Wenqiao Zhang, Sijing Li +12

    cs.CVcs.AIarXiv:2502.09838v32025
  60. Second-order Non-local Attention Networks for Person Re-identification

    Bryan, Xia, Yuan Gong +2

    cs.CVcs.AIcs.LGarXiv:1909.00295v12019