Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

421 to 480 of 18,830

  1. DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model

    Xueyi Liu, He Wang, Li Yi

    cs.ROcs.CVarXiv:2510.08556v12025
  2. Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation

    Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov +6

    cs.CVarXiv:2609.16874v12026
  3. VChain: Chain-of-Visual-Thought for Reasoning in Video Generation

    Ziqi Huang, Ning Yu, Gordon Chen +3

    cs.CVarXiv:2510.05094v22025
  4. GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

    Jiahe Li, Jiawei Zhang, Youmin Zhang +4

    cs.CVarXiv:2509.18090v22025
  5. Deep Convolutional Neural Networks for Interpretable Analysis of EEG Sleep Stage Scoring

    Albert Vilamala, Kristoffer H. Madsen, Lars K. Hansen

    cs.CVstat.MLarXiv:1710.00633v12017
  6. Deep Learning Based Steel Pipe Weld Defect Detection

    Dingming Yang, Yanrong Cui, Zeyu Yu +1

    cs.CVcs.AIarXiv:2104.14907v22021
  7. Motion Mamba: Efficient and Long Sequence Motion Generation

    Zeyu Zhang, Akide Liu, Ian Reid +3

    cs.CVarXiv:2403.07487v42024
  8. PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark

    Qiao Liu, Zhenyu He, Xin Li +1

    cs.CVarXiv:1801.05944v32018
  9. Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization

    Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3

    cs.CVcs.IRarXiv:2008.01403v22020
  10. FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

    Guangyu Sun, Shlok Kumar Mishra, Wentao Bao +8

    cs.CVarXiv:2609.16591v12026
  11. GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF

    Qiyu Dai, Yan Zhu, Yiran Geng +3

    cs.ROcs.CVarXiv:2210.06575v32022
  12. AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

    Jiaming Tan, Mingliang Zhai, Zhen Li +3

    cs.CVarXiv:2609.14462v12026
  13. Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

    Niantong Li, Guangzheng Hu, Weixu Qiao +35

    cs.CVarXiv:2605.28091v22026
  14. Frequency Perception Network for Camouflaged Object Detection

    Runmin Cong, Mengyao Sun, Sanyi Zhang +3

    cs.CVarXiv:2308.08924v22023
  15. Towards Zero-Shot Scale-Aware Monocular Depth Estimation

    Vitor Guizilini, Igor Vasiljevic, Dian Chen +2

    cs.CVcs.LGarXiv:2306.17253v12023
  16. PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

    Chuhao Chen, Peter Wonka, Chaoyang Wang +4

    cs.CVcs.AIcs.GRarXiv:2609.17521v12026
  17. BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

    Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14

    cs.CVarXiv:2609.15478v12026
  18. Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation

    Yin Wang, Mu Li, Jiapeng Liu +4

    cs.CVarXiv:2502.05534v12025
  19. KaiNinja: Extending Native 3D Generators to the Part Level

    Ruihan Yu, Lian Fu, Muyao Niu +9

    cs.GRcs.AIcs.CVarXiv:2609.15659v22026
  20. LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows

    Xiaofeng Mao, Peijia Lin, Shaohao Rui +3

    cs.CVarXiv:2609.15863v12026
  21. CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training

    Kihyun You, Jawook Gu, Jiyeon Ham +5

    cs.CVcs.LGarXiv:2310.13292v12023
  22. TJ4DRadSet: A 4D Radar Dataset for Autonomous Driving

    Lianqing Zheng, Zhixiong Ma, Xichan Zhu +9

    cs.CVcs.AIarXiv:2204.13483v32022
  23. RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

    Sibo Zhu, Shicheng Fan, Xinyue Wang +3

    cs.AIcs.CLcs.CVarXiv:2609.15364v12026
  24. EventVAD: Training-Free Event-Aware Video Anomaly Detection

    Yihua Shao, Haojin He, Sijie Li +11

    cs.CVarXiv:2504.13092v32025
  25. PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

    DeepCybo Team, Yu Bin, Haipeng Cao +51

    cs.CVcs.ROarXiv:2609.14973v12026
  26. Is Artificial Intelligence Generated Image Detection a Solved Problem?

    Ziqiang Li, Jiazhen Yan, Ziwen He +4

    cs.CVcs.CRarXiv:2505.12335v22025
  27. Generalizable Patch-Based Neural Rendering

    Mohammed Suhail, Carlos Esteves, Leonid Sigal +1

    cs.CVarXiv:2207.10662v22022
  28. Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs

    Yi Zhang, Bolin Ni, Xin-Sheng Chen +7

    cs.CVcs.AIarXiv:2510.13795v42025
  29. Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

    Biwen Lei, Yang Li, Xinhai Liu +97

    cs.CVarXiv:2509.12815v12025
  30. SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image

    Yu-Rou Tuan, Hao-Tang Tsui, Nicolas Ugrinovic +2

    cs.GRcs.CVarXiv:2609.13146v12026
  31. AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

    Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer +4

    cs.CLcs.AIcs.CVarXiv:2402.05602v22024
  32. 3D Aware Region Prompted Vision Language Model

    An-Chieh Cheng, Yang Fu, Yukang Chen +10

    cs.CVarXiv:2509.13317v12025
  33. Task-Embedded Control Networks for Few-Shot Imitation Learning

    Stephen James, Michael Bloesch, Andrew J. Davison

    cs.ROcs.AIcs.CVarXiv:1810.03237v12018
  34. Inverse Path Tracing for Joint Material and Lighting Estimation

    Dejan Azinović, Tzu-Mao Li, Anton Kaplanyan +1

    cs.CVarXiv:1903.07145v12019
  35. V-Thinker: Interactive Thinking with Images

    Runqi Qiao, Qiuna Tan, Minghan Yang +11

    cs.CVarXiv:2511.04460v22025
  36. Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

    Kai Zou, Ziqi Huang, Yuhao Dong +7

    cs.CVarXiv:2510.13759v32025
  37. Text2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors

    Tayeba Qazi, Brejesh Lall, Prerana Mukherjee

    cs.CVarXiv:2609.03585v22026
  38. Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems

    Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory +2

    cs.CVcs.CRarXiv:2609.03453v12026
  39. PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding

    Penghao Wang, Yiyang He, Xin Lv +4

    cs.CVarXiv:2510.20155v32025
  40. MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

    Jinkun Hao, Naifu Liang, Zhen Luo +8

    cs.CVcs.ROarXiv:2509.22281v12025
  41. The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding

    Weichen Fan, Haiwen Diao, Quan Wang +2

    cs.CVarXiv:2512.19693v52025
  42. Exploring Unlabeled Faces for Novel Attribute Discovery

    Hyojin Bahng, Sunghyo Chung, Seungjoo Yoo +1

    cs.CVarXiv:1912.03085v12019
  43. AnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head Reenactment

    Kangyeol Kim, Sunghyun Park, Jaeseong Lee +3

    cs.AIcs.CVarXiv:2111.07640v22021
  44. Improving Scene Text Recognition for Character-Level Long-Tailed Distribution

    Sunghyun Park, Sunghyo Chung, Jungsoo Lee +1

    cs.CVcs.CLarXiv:2304.08592v12023
  45. Test-Time Logit Prompting for Source-Free Missing Modality Adaptation

    Taixi Chen, Nancy Guo

    cs.CVarXiv:2609.02039v12026
  46. Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning

    Jiyoung Park, InJae Oh, Jung Uk Kim

    cs.CVarXiv:2609.01136v12026
  47. BenthicFlow: Generating Extensible Underwater Environments via Flow Matching

    Joaquín Figueira, Camile Lendering, Manfred Gonzalez-Hernandez +3

    cs.CVarXiv:2608.23173v12026
  48. RiT: Vanilla Diffusion Transformers Suffice in Representation Space

    Le Zhang, Ning Mang, Aishwarya Agrawal

    cs.CVarXiv:2605.21981v12026
  49. TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

    Jisu Nam, Jahyeok Koo, Soowon Son +4

    cs.CVarXiv:2605.12587v12026
  50. ELT: Elastic Looped Transformers for Visual Generation

    Sahil Goyal, Swayam Agrawal, Gautham Govind Anil +3

    cs.CVarXiv:2604.09168v32026
  51. Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs

    Nimrod Shabtay, Moshe Kimhi, Artem Spector +5

    cs.CVcs.AIarXiv:2603.16932v12026
  52. Panoramic Affordance Prediction

    Zixin Zhang, Chenfei Liao, Hongfei Zhang +10

    cs.CVcs.ROarXiv:2603.15558v12026
  53. LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency

    Weilong Yan, Haipeng Li, Hao Xu +4

    cs.CVcs.ROarXiv:2602.18735v22026
  54. EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models

    Dingkun Liu, Yuheng Chen, Zhu Chen +5

    cs.LGcs.CVarXiv:2601.17883v32026
  55. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

    Yi-Fan Zhang, Huanyu Zhang, Haochen Tian +10

    cs.CVarXiv:2408.13257v32024
  56. 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

    Guanjun Wu, Taoran Yi, Jiemin Fang +6

    cs.CVcs.GRarXiv:2310.08528v32023
  57. TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement

    Carl Doersch, Yi Yang, Mel Vecerik +5

    cs.CVarXiv:2306.08637v22023
  58. Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM

    Hengyi Wang, Jingwen Wang, Lourdes Agapito

    cs.CVarXiv:2304.14377v12023
  59. Symbolic Discovery of Optimization Algorithms

    Xiangning Chen, Chen Liang, Da Huang +9

    cs.LGcs.AIcs.CLarXiv:2302.06675v42023
  60. CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection

    Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki +1

    cs.CVarXiv:2212.05136v32022