Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,121 to 6,180 of 18,866
AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge
Vivek Chavan, Jörg Krüger
cs.CVcs.AIcs.HCarXiv:2608.22617v12026Metric and Kernel Learning using a Linear Transformation
Prateek Jain, Brian Kulis, Jason V. Davis +1
cs.LGcs.CVcs.IRarXiv:0910.5932v12009MaskedFace-Net -- A Dataset of Correctly/Incorrectly Masked Face Images in the Context of COVID-19
Adnane Cabani, Karim Hammoudi, Halim Benhabiles +1
cs.CVeess.IVarXiv:2008.08016v12020DeepSphere: Efficient spherical Convolutional Neural Network with HEALPix sampling for cosmological applications
Nathanaël Perraudin, Michaël Defferrard, Tomasz Kacprzak +1
astro-ph.COastro-ph.IMcs.AIarXiv:1810.12186v22018UniTok: A Unified Tokenizer for Visual Generation and Understanding
Chuofan Ma, Yi Jiang, Junfeng Wu +5
cs.CVcs.AIarXiv:2502.20321v32025Locally Attentional SDF Diffusion for Controllable 3D Shape Generation
Xin-Yang Zheng, Hao Pan, Peng-Shuai Wang +3
cs.CVcs.GRarXiv:2305.04461v22023Beyond Deep Residual Learning for Image Restoration: Persistent Homology-Guided Manifold Simplification
Woong Bae, Jaejun Yoo, Jong Chul Ye
cs.CVarXiv:1611.06345v42016FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
Shangzhan Zhang, Jianyuan Wang, Yinghao Xu +5
cs.CVarXiv:2502.12138v72025VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
Kai Deng, Zexin Ti, Jiawei Xu +2
cs.CVarXiv:2507.16443v22025CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling
Xin Shen, Chengyou Jia, Keshuo Xing +6
cs.CVarXiv:2609.01479v12026RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching
Charles Corbière, Léo Machado, Aubin Charley +3
cs.CVarXiv:2609.01470v12026AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
Kwan Yun, Serin Yoon, Sunjin Jung +3
cs.GRcs.CVcs.MMarXiv:2608.16143v12026TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
Dong Chen, Kenneth M. C. Cheung
cs.CVcs.AIcs.LGarXiv:2608.16122v12026A survey of AI-generated voices and their detection
Chengzhe Sun, Tianle Yang, Siwei Lyu
cs.AIcs.CVarXiv:2608.15411v12026Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
Thomas Manzini, Priyankari Perali, Raisa Karnik +2
cs.CVcs.AIarXiv:2608.14942v12026Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils
Karel Becerra, Boris Mederos, Dean Snow +1
cs.CVcs.AIcs.LGarXiv:2608.14539v12026Summaries:한국어Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Zian Meng, Zhen Li, Chuanhao Li +2
cs.CVcs.AIarXiv:2608.14530v12026PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
Yuyang Liu, Yanqing Shen, Ruike Chen +22
cs.ROcs.CVarXiv:2608.14284v12026Summaries:한국어Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning
Zixuan Wang, Yixin Hu, Wen Li +4
cs.CVarXiv:2609.00656v12026Seedream 3.0 Technical Report
Yu Gao, Lixue Gong, Qiushan Guo +28
cs.CVarXiv:2504.11346v32025Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
Zekai Gu, Rui Yan, Jiahao Lu +9
cs.CVcs.AIcs.GRarXiv:2501.03847v22025AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Yaxin Luo, Haobin Jiang, Jialv Zou +11
cs.CVcs.AIcs.CLarXiv:2608.13560v12026Summaries:한국어Unified Video Action Model
Shuang Li, Yihuai Gao, Dorsa Sadigh +1
cs.ROcs.CVarXiv:2503.00200v32025Learning elementary structures for 3D shape generation and matching
Theo Deprelle, Thibault Groueix, Matthew Fisher +3
cs.CVcs.AIarXiv:1908.04725v22019HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Dairu Liu, Zekun Qi, Jiayu Zeng +11
cs.ROcs.AIcs.CVarXiv:2608.13555v12026CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers
Ebenezer Tarubinga
cs.CVcs.LGeess.IVarXiv:2608.12773v12026SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Enze Xie, Junsong Chen, Yuyang Zhao +11
cs.CVarXiv:2501.18427v42025AVA-Encoder: Towards Agent-Native Video Representation Learning
Chuyue Li, Jinpeng Yu, Haozhe Wang +7
cs.CVcs.CLarXiv:2608.12313v12026Hierarchical Boundary-Aware Neural Encoder for Video Captioning
Lorenzo Baraldi, Costantino Grana, Rita Cucchiara
cs.CVarXiv:1611.09312v32016Blockwisely Supervised Neural Architecture Search with Knowledge Distillation
Changlin Li, Jiefeng Peng, Liuchun Yuan +4
cs.CVcs.LGcs.NEarXiv:1911.13053v22019Stepwise Goal-Driven Networks for Trajectory Prediction
Chuhua Wang, Yuchen Wang, Mingze Xu +1
cs.CVarXiv:2103.14107v32021HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes
Zan Wang, Yixin Chen, Tengyu Liu +3
cs.CVcs.AIarXiv:2210.09729v12022Weakly-supervised learning of visual relations
Julia Peyre, Ivan Laptev, Cordelia Schmid +1
cs.CVarXiv:1707.09472v12017FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
Yuren Cong, Mengmeng Xu, Christian Simon +7
cs.CVarXiv:2310.05922v32023Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
Zhiyuan You, Xin Cai, Jinjin Gu +2
cs.CVarXiv:2501.11561v32025From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Zepeng Wang, Jiagao Hu, Fuhao Li +3
cs.CVcs.AIeess.IVarXiv:2608.11562v12026Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae +1
cs.CVarXiv:2608.10708v12026InSight-doc: Agentic Visual Perception for Long-Document Understanding
Kaican Li, Weiyan Xie, Lewei Yao +4
cs.CVcs.CLcs.LGarXiv:2608.10628v12026End-to-End Lane Marker Detection via Row-wise Classification
Seungwoo Yoo, Heeseok Lee, Heesoo Myeong +4
cs.CVcs.LGarXiv:2005.08630v12020StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
Meng Wei, Chenyang Wan, Xiqian Yu +9
cs.ROcs.CVarXiv:2507.05240v22025SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models
Yuzhou Huang, Liangbin Xie, Xintao Wang +8
cs.CVarXiv:2312.06739v12023SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Yue Zhang, Yingzhao Jian, Yunqiu Xu +2
cs.CVarXiv:2608.05137v32026LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Ziyu Ma, Hailang Huang, Shun Zou +5
cs.CVarXiv:2608.01964v12026Spectral Prior for Reducing Exposure Bias in Diffusion Models
Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
cs.CVarXiv:2607.22091v12026ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Xiaozhong Lyu, Gen Li, Zhiyin Qian +3
cs.CVcs.AIarXiv:2607.17790v120262018 Robotic Scene Segmentation Challenge
Max Allan, Satoshi Kondo, Sebastian Bodenstedt +38
cs.CVcs.ROarXiv:2001.11190v32020Unified Reward Model for Multimodal Understanding and Generation
Yibin Wang, Yuhang Zang, Hao Li +2
cs.CVarXiv:2503.05236v22025Multiplayer Interactive World Models with Representation Autoencoders
Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary +24
cs.CVcs.AIcs.LGarXiv:2607.05352v22026TESSERA v2: Scaling Pixel-wise Earth Foundation Models
Zhengpeng Feng, Sadiq Jaffer, Ira Shokar +12
cs.CVcs.LGarXiv:2607.03949v22026Learning RoI Transformer for Detecting Oriented Objects in Aerial Images
Jian Ding, Nan Xue, Yang Long +2
cs.CVarXiv:1812.00155v12018SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
Paul Engstler, Iro Laina, Christian Rupprecht +1
cs.CVarXiv:2607.05392v12026VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
Wenzhuo Xu, Yuchen Zhu, Chongjian Ge +8
cs.CVarXiv:2609.03153v12026Attending to Multimodal Generation One Token at a Time
Varun Gupta, Vineet Gandhi, Makarand Tapaswi
cs.CVcs.AIarXiv:2607.03738v12026OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
Donghyun Lee, Jitesh Chavan, Duy Nguyen +5
cs.CVcs.AIcs.LGarXiv:2607.02461v12026AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
Haiyang Li, Yuming Fu, Qun Song +4
cs.CVarXiv:2607.02271v12026Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
Junha Jung, Minbyul Jeong, Suhyeon Lim +5
cs.CVcs.AIarXiv:2606.31825v12026Summaries:한국어TerraDiT-$Ω$: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive
Brian Wei, Srikumar Sastry, Daniel Cher +2
cs.CVarXiv:2606.31029v12026Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
Zhiqi Li, Chengrui Dong, Zhenhua Du +6
cs.CVarXiv:2606.30045v12026Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
Jongoh Jeong, Sun-Kyung Lee, Kuk-Jin Yoon
cs.CVcs.AIarXiv:2606.29464v12026PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Peiwen Zhang, Yufan Deng, Shangkun Sun +11
cs.CVcs.AIcs.ROarXiv:2606.28128v12026Summaries:한국어