Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,781 to 15,840 of 18,866
Few-shot Object Detection via Feature Reweighting
Bingyi Kang, Zhuang Liu, Xin Wang +3
cs.CVarXiv:1812.01866v22018MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
Weixiang Shen, Chengzhi Shen, Yanzhu Hu +12
cs.CVarXiv:2603.24649v22026CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin +5
cs.LGcs.AIcs.CVarXiv:2603.24440v12026Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
Mohit Shridhar, Lucas Manuelli, Dieter Fox
cs.ROcs.AIcs.CLarXiv:2209.05451v22022ShowUI-Aloha: Human-Taught GUI Agent
Yichun Zhang, Xiangwu Guo, Yauhong Goh +5
cs.CVarXiv:2601.07181v12026Siamese Box Adaptive Network for Visual Tracking
Zedu Chen, Bineng Zhong, Guorong Li +2
cs.CVarXiv:2003.06761v22020TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
Yu Xu, Hongbin Yan, Juan Cao +11
cs.CVcs.AIarXiv:2601.08881v22026DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction
Jiayang Shi, Daniel M. Pelt, K. Joost Batenburg
eess.IVcs.AIcs.CVarXiv:2602.18589v12026SAPIEN: A SimulAted Part-based Interactive ENvironment
Fanbo Xiang, Yuzhe Qin, Kaichun Mo +11
cs.CVcs.ROarXiv:2003.08515v12020MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
Zhang Li, Zhibo Lin, Qiang Liu +7
cs.CVcs.AIarXiv:2603.28130v12026How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai +3
cs.CVcs.AIcs.LGarXiv:2106.10270v22021MolmoPoint: Better Pointing for VLMs with Grounding Tokens
Christopher Clark, Yue Yang, Jae Sung Park +8
cs.CVcs.AIarXiv:2603.28069v12026The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli +2
cs.CVcs.AIcs.CLarXiv:1904.12584v12019AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation
Zhaohe Liao, Kaixun Jiang, Zhihang Liu +11
cs.CVarXiv:2603.28068v22026LLVIP: A Visible-infrared Paired Dataset for Low-light Vision
Xinyu Jia, Chuang Zhu, Minzhen Li +3
cs.CVcs.AIarXiv:2108.10831v42021Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan +4
cs.CVarXiv:2303.13439v12023EchoWM: Open and Enterable Omnimodal World Models
Songchun Zhang, Yaowei Li, Junhao Zhuang +19
cs.CVarXiv:2608.23189v12026Deep Image Prior
Dmitry Ulyanov, Andrea Vedaldi, Victor Lempitsky
cs.CVstat.MLarXiv:1711.10925v42017Semantic Understanding of Scenes through the ADE20K Dataset
Bolei Zhou, Hang Zhao, Xavier Puig +4
cs.CVarXiv:1608.05442v22016DiffusionBench: On Holistic Evaluation of Diffusion Transformers
Xingjian Leng, Jaskirat Singh, Zhanhao Liang +5
cs.CVarXiv:2606.24888v12026Deeply learned face representations are sparse, selective, and robust
Yi Sun, Xiaogang Wang, Xiaoou Tang
cs.CVarXiv:1412.1265v12014Enriching ImageNet with Human Similarity Judgments and Psychological Embeddings
Brett D. Roads, Bradley C. Love
cs.CVcs.LGarXiv:2011.11015v12020Summaries:한국어Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Michael M. Bronstein, Joan Bruna, Taco Cohen +1
cs.LGcs.AIcs.CGarXiv:2104.13478v22021Neural Motifs: Scene Graph Parsing with Global Context
Rowan Zellers, Mark Yatskar, Sam Thomson +1
cs.CVarXiv:1711.06640v22017EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
Wuyang Li, Yang Gao, Mariam Hassan +4
cs.CVcs.AIarXiv:2605.15042v12026DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
Kailai Feng, Yuxiang Wei, Bo Chen +5
cs.CVarXiv:2603.28713v12026MixFormer: End-to-End Tracking with Iterative Mixed Attention
Yutao Cui, Cheng Jiang, Limin Wang +1
cs.CVarXiv:2203.11082v22022A New Representation of Skeleton Sequences for 3D Action Recognition
Qiuhong Ke, Mohammed Bennamoun, Senjian An +2
cs.CVarXiv:1703.03492v32017M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network
Qijie Zhao, Tao Sheng, Yongtao Wang +4
cs.CVarXiv:1811.04533v32018MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking
Laura Leal-Taixé, Anton Milan, Ian Reid +2
cs.CVarXiv:1504.01942v12015F3Net: Fusion, Feedback and Focus for Salient Object Detection
Jun Wei, Shuhui Wang, Qingming Huang
cs.CVarXiv:1911.11445v12019Real-time Scene Text Detection with Differentiable Binarization
Minghui Liao, Zhaoyi Wan, Cong Yao +2
cs.CVarXiv:1911.08947v22019MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
Jiachun Li, Shaoping Huang, Zhuoran Jin +5
cs.CLcs.AIcs.CVarXiv:2603.02024v12026Improved Texture Networks: Maximizing Quality and Diversity in Feed-forward Stylization and Texture Synthesis
Dmitry Ulyanov, Andrea Vedaldi, Victor Lempitsky
cs.CVarXiv:1701.02096v22017AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
Mingyang Song, Haoyu Sun, Jiawei Gu +4
cs.AIcs.CLcs.CVarXiv:2601.18631v22026Learning Texture Transformer Network for Image Super-Resolution
Fuzhi Yang, Huan Yang, Jianlong Fu +2
cs.CVarXiv:2006.04139v22020DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
Chenlong Deng, Mengjie Deng, Junjie Wu +10
cs.CVcs.IRarXiv:2602.10809v22026BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth
Mahdi Rad, Vincent Lepetit
cs.CVarXiv:1703.10896v22017ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
Changti Wu, Jiahuai Mao, Yuzhuo Miao +6
cs.CVcs.AIarXiv:2602.11636v12026Continual GUI Agents
Ziwei Liu, Borui Kang, Hangjie Yuan +4
cs.LGcs.CVarXiv:2601.20732v42026Phi-4-reasoning-vision-15B Technical Report
Jyoti Aneja, Michael Harrison, Neel Joshi +3
cs.AIcs.CVarXiv:2603.03975v12026Point Transformer V2: Grouped Vector Attention and Partition-based Pooling
Xiaoyang Wu, Yixing Lao, Li Jiang +2
cs.CVarXiv:2210.05666v22022NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture Search
Xuanyi Dong, Yi Yang
cs.CVarXiv:2001.00326v22020Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach
Yunchao Wei, Jiashi Feng, Xiaodan Liang +3
cs.CVarXiv:1703.08448v32017LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Benno Krojer, Shravan Nayak, Oscar Mañas +4
cs.CVcs.AIarXiv:2602.00462v52026Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC)
David Gutman, Noel C. F. Codella, Emre Celebi +4
cs.CVarXiv:1605.01397v12016DeepOrgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation
Holger R. Roth, Le Lu, Amal Farag +4
cs.CVarXiv:1506.06448v12015Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
Yuanyuan Gao, Hao Li, Yifei Liu +14
cs.CVarXiv:2603.07660v12026Learning to Enhance Low-Light Image via Zero-Reference Deep Curve Estimation
Chongyi Li, Chunle Guo, Chen Change Loy
cs.CVarXiv:2103.00860v12021Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4
cs.CLcs.CVarXiv:2603.09095v32026PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization
Shunsuke Saito, Tomas Simon, Jason Saragih +1
cs.CVcs.GRarXiv:2004.00452v12020CondenseNet: An Efficient DenseNet using Learned Group Convolutions
Gao Huang, Shichen Liu, Laurens van der Maaten +1
cs.CVarXiv:1711.09224v22017Towards Accurate Multi-person Pose Estimation in the Wild
George Papandreou, Tyler Zhu, Nori Kanazawa +4
cs.CVarXiv:1701.01779v22017PETR: Position Embedding Transformation for Multi-View 3D Object Detection
Yingfei Liu, Tiancai Wang, Xiangyu Zhang +1
cs.CVarXiv:2203.05625v32022Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
Haotian Tang, Zhijian Liu, Shengyu Zhao +4
cs.CVarXiv:2007.16100v22020SpiderCNN: Deep Learning on Point Sets with Parameterized Convolutional Filters
Yifan Xu, Tianqi Fan, Mingye Xu +2
cs.CVarXiv:1803.11527v32018A Large-Scale Car Dataset for Fine-Grained Categorization and Verification
Linjie Yang, Ping Luo, Chen Change Loy +1
cs.CVcs.AIarXiv:1506.08959v22015PixelGen: Improving Pixel Diffusion with Perceptual Supervision
Zehong Ma, Ruihan Xu, Shiliang Zhang
cs.CVcs.AIarXiv:2602.02493v22026FILIP: Fine-grained Interactive Language-Image Pre-Training
Lewei Yao, Runhui Huang, Lu Hou +7
cs.CVcs.LGarXiv:2111.07783v12021SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models
Jiwoo Chung, Sangeek Hyun, MinKyu Lee +5
cs.CVarXiv:2602.18993v22026