Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
421 to 480 of 18,830
DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model
Xueyi Liu, He Wang, Li Yi
cs.ROcs.CVarXiv:2510.08556v12025Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation
Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov +6
cs.CVarXiv:2609.16874v12026VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
Ziqi Huang, Ning Yu, Gordon Chen +3
cs.CVarXiv:2510.05094v22025GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
Jiahe Li, Jiawei Zhang, Youmin Zhang +4
cs.CVarXiv:2509.18090v22025Deep Convolutional Neural Networks for Interpretable Analysis of EEG Sleep Stage Scoring
Albert Vilamala, Kristoffer H. Madsen, Lars K. Hansen
cs.CVstat.MLarXiv:1710.00633v12017Deep Learning Based Steel Pipe Weld Defect Detection
Dingming Yang, Yanrong Cui, Zeyu Yu +1
cs.CVcs.AIarXiv:2104.14907v22021Motion Mamba: Efficient and Long Sequence Motion Generation
Zeyu Zhang, Akide Liu, Ian Reid +3
cs.CVarXiv:2403.07487v42024PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark
Qiao Liu, Zhenyu He, Xin Li +1
cs.CVarXiv:1801.05944v32018Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3
cs.CVcs.IRarXiv:2008.01403v22020FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Guangyu Sun, Shlok Kumar Mishra, Wentao Bao +8
cs.CVarXiv:2609.16591v12026GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF
Qiyu Dai, Yan Zhu, Yiran Geng +3
cs.ROcs.CVarXiv:2210.06575v32022AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Jiaming Tan, Mingliang Zhai, Zhen Li +3
cs.CVarXiv:2609.14462v12026Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Niantong Li, Guangzheng Hu, Weixu Qiao +35
cs.CVarXiv:2605.28091v22026Frequency Perception Network for Camouflaged Object Detection
Runmin Cong, Mengyao Sun, Sanyi Zhang +3
cs.CVarXiv:2308.08924v22023Towards Zero-Shot Scale-Aware Monocular Depth Estimation
Vitor Guizilini, Igor Vasiljevic, Dian Chen +2
cs.CVcs.LGarXiv:2306.17253v12023PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
Chuhao Chen, Peter Wonka, Chaoyang Wang +4
cs.CVcs.AIcs.GRarXiv:2609.17521v12026BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14
cs.CVarXiv:2609.15478v12026Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
Yin Wang, Mu Li, Jiapeng Liu +4
cs.CVarXiv:2502.05534v12025KaiNinja: Extending Native 3D Generators to the Part Level
Ruihan Yu, Lian Fu, Muyao Niu +9
cs.GRcs.AIcs.CVarXiv:2609.15659v22026LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Xiaofeng Mao, Peijia Lin, Shaohao Rui +3
cs.CVarXiv:2609.15863v12026CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training
Kihyun You, Jawook Gu, Jiyeon Ham +5
cs.CVcs.LGarXiv:2310.13292v12023TJ4DRadSet: A 4D Radar Dataset for Autonomous Driving
Lianqing Zheng, Zhixiong Ma, Xichan Zhu +9
cs.CVcs.AIarXiv:2204.13483v32022RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Sibo Zhu, Shicheng Fan, Xinyue Wang +3
cs.AIcs.CLcs.CVarXiv:2609.15364v12026EventVAD: Training-Free Event-Aware Video Anomaly Detection
Yihua Shao, Haojin He, Sijie Li +11
cs.CVarXiv:2504.13092v32025PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
DeepCybo Team, Yu Bin, Haipeng Cao +51
cs.CVcs.ROarXiv:2609.14973v12026Is Artificial Intelligence Generated Image Detection a Solved Problem?
Ziqiang Li, Jiazhen Yan, Ziwen He +4
cs.CVcs.CRarXiv:2505.12335v22025Generalizable Patch-Based Neural Rendering
Mohammed Suhail, Carlos Esteves, Leonid Sigal +1
cs.CVarXiv:2207.10662v22022Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
Yi Zhang, Bolin Ni, Xin-Sheng Chen +7
cs.CVcs.AIarXiv:2510.13795v42025Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation
Biwen Lei, Yang Li, Xinhai Liu +97
cs.CVarXiv:2509.12815v12025SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image
Yu-Rou Tuan, Hao-Tang Tsui, Nicolas Ugrinovic +2
cs.GRcs.CVarXiv:2609.13146v12026AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer +4
cs.CLcs.AIcs.CVarXiv:2402.05602v220243D Aware Region Prompted Vision Language Model
An-Chieh Cheng, Yang Fu, Yukang Chen +10
cs.CVarXiv:2509.13317v12025Task-Embedded Control Networks for Few-Shot Imitation Learning
Stephen James, Michael Bloesch, Andrew J. Davison
cs.ROcs.AIcs.CVarXiv:1810.03237v12018Inverse Path Tracing for Joint Material and Lighting Estimation
Dejan Azinović, Tzu-Mao Li, Anton Kaplanyan +1
cs.CVarXiv:1903.07145v12019V-Thinker: Interactive Thinking with Images
Runqi Qiao, Qiuna Tan, Minghan Yang +11
cs.CVarXiv:2511.04460v22025Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
Kai Zou, Ziqi Huang, Yuhao Dong +7
cs.CVarXiv:2510.13759v32025Text2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors
Tayeba Qazi, Brejesh Lall, Prerana Mukherjee
cs.CVarXiv:2609.03585v22026Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems
Jannatul Masruk Mukta, Rifa Sanjida, Adrita Rahman Tory +2
cs.CVcs.CRarXiv:2609.03453v12026PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
Penghao Wang, Yiyang He, Xin Lv +4
cs.CVarXiv:2510.20155v32025MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
Jinkun Hao, Naifu Liang, Zhen Luo +8
cs.CVcs.ROarXiv:2509.22281v12025The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
Weichen Fan, Haiwen Diao, Quan Wang +2
cs.CVarXiv:2512.19693v52025Exploring Unlabeled Faces for Novel Attribute Discovery
Hyojin Bahng, Sunghyo Chung, Seungjoo Yoo +1
cs.CVarXiv:1912.03085v12019AnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head Reenactment
Kangyeol Kim, Sunghyun Park, Jaeseong Lee +3
cs.AIcs.CVarXiv:2111.07640v22021Improving Scene Text Recognition for Character-Level Long-Tailed Distribution
Sunghyun Park, Sunghyo Chung, Jungsoo Lee +1
cs.CVcs.CLarXiv:2304.08592v12023Test-Time Logit Prompting for Source-Free Missing Modality Adaptation
Taixi Chen, Nancy Guo
cs.CVarXiv:2609.02039v12026Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
Jiyoung Park, InJae Oh, Jung Uk Kim
cs.CVarXiv:2609.01136v12026BenthicFlow: Generating Extensible Underwater Environments via Flow Matching
Joaquín Figueira, Camile Lendering, Manfred Gonzalez-Hernandez +3
cs.CVarXiv:2608.23173v12026RiT: Vanilla Diffusion Transformers Suffice in Representation Space
Le Zhang, Ning Mang, Aishwarya Agrawal
cs.CVarXiv:2605.21981v12026TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
Jisu Nam, Jahyeok Koo, Soowon Son +4
cs.CVarXiv:2605.12587v12026ELT: Elastic Looped Transformers for Visual Generation
Sahil Goyal, Swayam Agrawal, Gautham Govind Anil +3
cs.CVarXiv:2604.09168v32026Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
Nimrod Shabtay, Moshe Kimhi, Artem Spector +5
cs.CVcs.AIarXiv:2603.16932v12026Panoramic Affordance Prediction
Zixin Zhang, Chenfei Liao, Hongfei Zhang +10
cs.CVcs.ROarXiv:2603.15558v12026LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency
Weilong Yan, Haipeng Li, Hao Xu +4
cs.CVcs.ROarXiv:2602.18735v22026EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models
Dingkun Liu, Yuheng Chen, Zhu Chen +5
cs.LGcs.CVarXiv:2601.17883v32026MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Yi-Fan Zhang, Huanyu Zhang, Haochen Tian +10
cs.CVarXiv:2408.13257v320244D Gaussian Splatting for Real-Time Dynamic Scene Rendering
Guanjun Wu, Taoran Yi, Jiemin Fang +6
cs.CVcs.GRarXiv:2310.08528v32023TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement
Carl Doersch, Yi Yang, Mel Vecerik +5
cs.CVarXiv:2306.08637v22023Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM
Hengyi Wang, Jingwen Wang, Lourdes Agapito
cs.CVarXiv:2304.14377v12023Symbolic Discovery of Optimization Algorithms
Xiangning Chen, Chen Liang, Da Huang +9
cs.LGcs.AIcs.CLarXiv:2302.06675v42023CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki +1
cs.CVarXiv:2212.05136v32022