Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,901 to 15,960 of 18,830
Joint Detection and Identification Feature Learning for Person Search
Tong Xiao, Shuang Li, Bochao Wang +2
cs.CVarXiv:1604.01850v32016SRDiff: Single Image Super-Resolution with Diffusion Probabilistic Models
Haoying Li, Yifan Yang, Meng Chang +4
cs.CVarXiv:2104.14951v22021SAMTok: Representing Any Mask with Two Words
Yikang Zhou, Tao Zhang, Dengxian Gong +13
cs.CVarXiv:2601.16093v22026Asymmetric Loss For Multi-Label Classification
Emanuel Ben-Baruch, Tal Ridnik, Nadav Zamir +4
cs.CVcs.LGarXiv:2009.14119v42020Deep Audio-Visual Speech Recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior +2
cs.CVarXiv:1809.02108v22018DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning
Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi +8
cs.LGcs.CVarXiv:2204.04799v22022Demystifying Video Reasoning
Ruisi Wang, Zhongang Cai, Fanyi Pu +11
cs.CVcs.AIarXiv:2603.16870v32026ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation
Sachin Mehta, Mohammad Rastegari, Anat Caspi +2
cs.CVarXiv:1803.06815v32018Semantic Autoencoder for Zero-Shot Learning
Elyor Kodirov, Tao Xiang, Shaogang Gong
cs.CVarXiv:1704.08345v12017Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
Kaleb Newman, Tyler Zhu, Olga Russakovsky
cs.CVarXiv:2603.30043v12026Autoregressive Image Generation using Residual Quantization
Doyup Lee, Chiheon Kim, Saehoon Kim +2
cs.CVcs.LGarXiv:2203.01941v22022PhyRPR: Training-Free Physics-Constrained Video Generation
Yibo Zhao, Hengjia Li, Xiaofei He +1
cs.CVarXiv:2601.09255v12026MS-TCN: Multi-Stage Temporal Convolutional Network for Action Segmentation
Yazan Abu Farha, Juergen Gall
cs.CVarXiv:1903.01945v22019LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
Jiazheng Xing, Fei Du, Hangjie Yuan +7
cs.CVcs.AIarXiv:2603.20192v12026BARF: Bundle-Adjusting Neural Radiance Fields
Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba +1
cs.CVcs.GRcs.LGarXiv:2104.06405v22021DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation
Gwanghyun Kim, Taesung Kwon, Jong Chul Ye
cs.CVcs.AIcs.LGarXiv:2110.02711v62021FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Qian Chen, Jinlan Fu, Changsong Li +3
cs.CLcs.CVcs.MMarXiv:2601.13836v22026One-Shot Learning for Semantic Segmentation
Amirreza Shaban, Shray Bansal, Zhen Liu +2
cs.CVarXiv:1709.03410v12017RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment
Liyao Jiang, Ruichen Chen, Chao Gao +1
cs.CVcs.AIarXiv:2603.00483v12026Manifold-Aware Exploration for Reinforcement Learning in Video Generation
Mingzhe Zheng, Weijie Kong, Yue Wu +9
cs.CVcs.AIarXiv:2603.21872v12026Not-so-supervised: a survey of semi-supervised, multi-instance, and transfer learning in medical image analysis
Veronika Cheplygina, Marleen de Bruijne, Josien P. W. Pluim
cs.CVarXiv:1804.06353v22018PersonaVLM: Long-Term Personalized Multimodal LLMs
Chang Nie, Chaoyou Fu, Yifan Zhang +2
cs.CLcs.CVarXiv:2604.13074v12026SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
Chong Xia, Kai Zhu, Zizhuo Wang +3
cs.CVarXiv:2603.02133v22026Learning Pixel-level Semantic Affinity with Image-level Supervision for Weakly Supervised Semantic Segmentation
Jiwoon Ahn, Suha Kwak
cs.CVarXiv:1803.10464v22018Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis
Yucheng Tang, Dong Yang, Wenqi Li +5
cs.CVcs.AIcs.LGarXiv:2111.14791v22021PEARL: Personalized Streaming Video Understanding Model
Yuanhong Zheng, Ruichuan An, Xiaopeng Lin +10
cs.CVcs.AIcs.IRarXiv:2603.20422v12026Recurrent Squeeze-and-Excitation Context Aggregation Net for Single Image Deraining
Xia Li, Jianlong Wu, Zhouchen Lin +2
cs.CVarXiv:1807.05698v22018tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
Chen Wang, Hao Tan, Wang Yifan +6
cs.CVarXiv:2602.20160v22026HRank: Filter Pruning using High-Rank Feature Map
Mingbao Lin, Rongrong Ji, Yan Wang +4
cs.CVarXiv:2002.10179v22020Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
cs.LGcs.AIcs.CLarXiv:2601.21037v12026A Comprehensive Survey of Deep Learning for Image Captioning
Md. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin +1
cs.CVcs.LGstat.MLarXiv:1810.04020v22018CAD2RL: Real Single-Image Flight without a Single Real Image
Fereshteh Sadeghi, Sergey Levine
cs.LGcs.CVcs.ROarXiv:1611.04201v42016Deep Reconstruction-Classification Networks for Unsupervised Domain Adaptation
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang +2
cs.CVcs.AIcs.LGarXiv:1607.03516v22016iMAP: Implicit Mapping and Positioning in Real-Time
Edgar Sucar, Shikun Liu, Joseph Ortiz +1
cs.CVarXiv:2103.12352v22021Pose Guided Person Image Generation
Liqian Ma, Xu Jia, Qianru Sun +3
cs.CVarXiv:1705.09368v62017Dual Path Networks
Yunpeng Chen, Jianan Li, Huaxin Xiao +3
cs.CVarXiv:1707.01629v22017CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion
Zixiang Zhao, Haowen Bai, Jiangshe Zhang +5
cs.CVarXiv:2211.14461v22022Causal Motion Diffusion Models for Autoregressive Motion Generation
Qing Yu, Akihisa Watanabe, Kent Fujiwara
cs.CVarXiv:2602.22594v12026LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
Zebin You, Xiaolu Zhang, Jun Zhou +2
cs.CVcs.LGarXiv:2603.01068v12026AVControl: Efficient Framework for Training Audio-Visual Controls
Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem +6
cs.CVcs.MMcs.SDarXiv:2603.24793v12026Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
Jiahao Lu, Jiayi Xu, Wenbo Hu +5
cs.CVarXiv:2603.02573v22026UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction
Michael Oechsle, Songyou Peng, Andreas Geiger
cs.CVarXiv:2104.10078v22021Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Jiyuan Wang, Chunyu Lin, Lei Sun +8
cs.CVcs.AIarXiv:2603.03143v22026Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation
Huaizu Jiang, Deqing Sun, Varun Jampani +3
cs.CVarXiv:1712.00080v22017Chain of World: World Model Thinking in Latent Motion
Fuxiang Yang, Donglin Di, Lulu Tang +6
cs.CVcs.AIcs.ROarXiv:2603.03195v12026Learned Primal-dual Reconstruction
Jonas Adler, Ozan Öktem
math.OCcs.CVcs.NEarXiv:1707.06474v32017Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
Hai Zhang, Siqi Liang, Li Chen +5
cs.CVcs.ROarXiv:2602.05827v12026Learning Memory-guided Normality for Anomaly Detection
Hyunjong Park, Jongyoun Noh, Bumsub Ham
cs.CVarXiv:2003.13228v12020COVID-CT-Dataset: A CT Scan Dataset about COVID-19
Xingyi Yang, Xuehai He, Jinyu Zhao +3
cs.LGcs.CVeess.IVarXiv:2003.13865v32020Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
Jie Zhu, Yiyang Su, Xiaoming Liu
cs.CVarXiv:2601.06993v22026Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
Dongwon Kim, Gawon Seo, Jinsung Lee +2
cs.CVcs.AIcs.ROarXiv:2603.05438v12026Action-Conditional Video Prediction using Deep Networks in Atari Games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee +2
cs.LGcs.AIcs.CVarXiv:1507.08750v22015Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral +1
cs.CLcs.AIcs.CVarXiv:2104.08773v42021SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh Rendering
Antoine Guédon, Vincent Lepetit
cs.GRcs.CVarXiv:2311.12775v32023Geometry-Aware Rotary Position Embedding for Consistent Video World Model
Chendong Xiang, Jiajun Liu, Jintao Zhang +7
cs.CVarXiv:2602.07854v32026ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
Samin Mahdizadeh Sani, Max Ku, Nima Jamali +23
cs.GRcs.AIcs.CVarXiv:2603.27862v12026Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja +1
cs.CVarXiv:1710.01992v32017Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker +1
cs.LGcs.CVstat.MLarXiv:1506.07365v32015FlowAct-R1: Towards Interactive Humanoid Video Generation
Lizhen Wang, Yongming Zhu, Zhipeng Ge +15
cs.CVcs.AIarXiv:2601.10103v12026Decoupled Knowledge Distillation
Borui Zhao, Quan Cui, Renjie Song +2
cs.CVcs.AIarXiv:2203.08679v22022