Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

15,781 to 15,840 of 18,866

  1. Few-shot Object Detection via Feature Reweighting

    Bingyi Kang, Zhuang Liu, Xin Wang +3

    cs.CVarXiv:1812.01866v22018
  2. MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

    Weixiang Shen, Chengzhi Shen, Yanzhu Hu +12

    cs.CVarXiv:2603.24649v22026
  3. CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents

    Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin +5

    cs.LGcs.AIcs.CVarXiv:2603.24440v12026
  4. Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation

    Mohit Shridhar, Lucas Manuelli, Dieter Fox

    cs.ROcs.AIcs.CLarXiv:2209.05451v22022
  5. ShowUI-Aloha: Human-Taught GUI Agent

    Yichun Zhang, Xiangwu Guo, Yauhong Goh +5

    cs.CVarXiv:2601.07181v12026
  6. Siamese Box Adaptive Network for Visual Tracking

    Zedu Chen, Bineng Zhong, Guorong Li +2

    cs.CVarXiv:2003.06761v22020
  7. TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts

    Yu Xu, Hongbin Yan, Juan Cao +11

    cs.CVcs.AIarXiv:2601.08881v22026
  8. DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction

    Jiayang Shi, Daniel M. Pelt, K. Joost Batenburg

    eess.IVcs.AIcs.CVarXiv:2602.18589v12026
  9. SAPIEN: A SimulAted Part-based Interactive ENvironment

    Fanbo Xiang, Yuzhe Qin, Kaichun Mo +11

    cs.CVcs.ROarXiv:2003.08515v12020
  10. MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

    Zhang Li, Zhibo Lin, Qiang Liu +7

    cs.CVcs.AIarXiv:2603.28130v12026
  11. How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

    Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai +3

    cs.CVcs.AIcs.LGarXiv:2106.10270v22021
  12. MolmoPoint: Better Pointing for VLMs with Grounding Tokens

    Christopher Clark, Yue Yang, Jae Sung Park +8

    cs.CVcs.AIarXiv:2603.28069v12026
  13. The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

    Jiayuan Mao, Chuang Gan, Pushmeet Kohli +2

    cs.CVcs.AIcs.CLarXiv:1904.12584v12019
  14. AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation

    Zhaohe Liao, Kaixun Jiang, Zhihang Liu +11

    cs.CVarXiv:2603.28068v22026
  15. LLVIP: A Visible-infrared Paired Dataset for Low-light Vision

    Xinyu Jia, Chuang Zhu, Minzhen Li +3

    cs.CVcs.AIarXiv:2108.10831v42021
  16. Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

    Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan +4

    cs.CVarXiv:2303.13439v12023
  17. EchoWM: Open and Enterable Omnimodal World Models

    Songchun Zhang, Yaowei Li, Junhao Zhuang +19

    cs.CVarXiv:2608.23189v12026
  18. Deep Image Prior

    Dmitry Ulyanov, Andrea Vedaldi, Victor Lempitsky

    cs.CVstat.MLarXiv:1711.10925v42017
  19. Semantic Understanding of Scenes through the ADE20K Dataset

    Bolei Zhou, Hang Zhao, Xavier Puig +4

    cs.CVarXiv:1608.05442v22016
  20. DiffusionBench: On Holistic Evaluation of Diffusion Transformers

    Xingjian Leng, Jaskirat Singh, Zhanhao Liang +5

    cs.CVarXiv:2606.24888v12026
  21. Deeply learned face representations are sparse, selective, and robust

    Yi Sun, Xiaogang Wang, Xiaoou Tang

    cs.CVarXiv:1412.1265v12014
  22. Enriching ImageNet with Human Similarity Judgments and Psychological Embeddings

    Brett D. Roads, Bradley C. Love

    cs.CVcs.LGarXiv:2011.11015v12020
    Summaries:한국어
  23. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

    Michael M. Bronstein, Joan Bruna, Taco Cohen +1

    cs.LGcs.AIcs.CGarXiv:2104.13478v22021
  24. Neural Motifs: Scene Graph Parsing with Global Context

    Rowan Zellers, Mark Yatskar, Sam Thomson +1

    cs.CVarXiv:1711.06640v22017
  25. EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

    Wuyang Li, Yang Gao, Mariam Hassan +4

    cs.CVcs.AIarXiv:2605.15042v12026
  26. DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing

    Kailai Feng, Yuxiang Wei, Bo Chen +5

    cs.CVarXiv:2603.28713v12026
  27. MixFormer: End-to-End Tracking with Iterative Mixed Attention

    Yutao Cui, Cheng Jiang, Limin Wang +1

    cs.CVarXiv:2203.11082v22022
  28. A New Representation of Skeleton Sequences for 3D Action Recognition

    Qiuhong Ke, Mohammed Bennamoun, Senjian An +2

    cs.CVarXiv:1703.03492v32017
  29. M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network

    Qijie Zhao, Tao Sheng, Yongtao Wang +4

    cs.CVarXiv:1811.04533v32018
  30. MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking

    Laura Leal-Taixé, Anton Milan, Ian Reid +2

    cs.CVarXiv:1504.01942v12015
  31. F3Net: Fusion, Feedback and Focus for Salient Object Detection

    Jun Wei, Shuhui Wang, Qingming Huang

    cs.CVarXiv:1911.11445v12019
  32. Real-time Scene Text Detection with Differentiable Binarization

    Minghui Liao, Zhaoyi Wan, Cong Yao +2

    cs.CVarXiv:1911.08947v22019
  33. MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

    Jiachun Li, Shaoping Huang, Zhuoran Jin +5

    cs.CLcs.AIcs.CVarXiv:2603.02024v12026
  34. Improved Texture Networks: Maximizing Quality and Diversity in Feed-forward Stylization and Texture Synthesis

    Dmitry Ulyanov, Andrea Vedaldi, Victor Lempitsky

    cs.CVarXiv:1701.02096v22017
  35. AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning

    Mingyang Song, Haoyu Sun, Jiawei Gu +4

    cs.AIcs.CLcs.CVarXiv:2601.18631v22026
  36. Learning Texture Transformer Network for Image Super-Resolution

    Fuzhi Yang, Huan Yang, Jianlong Fu +2

    cs.CVarXiv:2006.04139v22020
  37. DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

    Chenlong Deng, Mengjie Deng, Junjie Wu +10

    cs.CVcs.IRarXiv:2602.10809v22026
  38. BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth

    Mahdi Rad, Vincent Lepetit

    cs.CVarXiv:1703.10896v22017
  39. ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning

    Changti Wu, Jiahuai Mao, Yuzhuo Miao +6

    cs.CVcs.AIarXiv:2602.11636v12026
  40. Continual GUI Agents

    Ziwei Liu, Borui Kang, Hangjie Yuan +4

    cs.LGcs.CVarXiv:2601.20732v42026
  41. Phi-4-reasoning-vision-15B Technical Report

    Jyoti Aneja, Michael Harrison, Neel Joshi +3

    cs.AIcs.CVarXiv:2603.03975v12026
  42. Point Transformer V2: Grouped Vector Attention and Partition-based Pooling

    Xiaoyang Wu, Yixing Lao, Li Jiang +2

    cs.CVarXiv:2210.05666v22022
  43. NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture Search

    Xuanyi Dong, Yi Yang

    cs.CVarXiv:2001.00326v22020
  44. Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach

    Yunchao Wei, Jiashi Feng, Xiaodan Liang +3

    cs.CVarXiv:1703.08448v32017
  45. LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

    Benno Krojer, Shravan Nayak, Oscar Mañas +4

    cs.CVcs.AIarXiv:2602.00462v52026
  46. Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC)

    David Gutman, Noel C. F. Codella, Emre Celebi +4

    cs.CVarXiv:1605.01397v12016
  47. DeepOrgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation

    Holger R. Roth, Le Lu, Amal Farag +4

    cs.CVarXiv:1506.06448v12015
  48. Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

    Yuanyuan Gao, Hao Li, Yifei Liu +14

    cs.CVarXiv:2603.07660v12026
  49. Learning to Enhance Low-Light Image via Zero-Reference Deep Curve Estimation

    Chongyi Li, Chunle Guo, Chen Change Loy

    cs.CVarXiv:2103.00860v12021
  50. Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs

    Kaiser Sun, Xiaochuang Yuan, Hongjun Liu +4

    cs.CLcs.CVarXiv:2603.09095v32026
  51. PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization

    Shunsuke Saito, Tomas Simon, Jason Saragih +1

    cs.CVcs.GRarXiv:2004.00452v12020
  52. CondenseNet: An Efficient DenseNet using Learned Group Convolutions

    Gao Huang, Shichen Liu, Laurens van der Maaten +1

    cs.CVarXiv:1711.09224v22017
  53. Towards Accurate Multi-person Pose Estimation in the Wild

    George Papandreou, Tyler Zhu, Nori Kanazawa +4

    cs.CVarXiv:1701.01779v22017
  54. PETR: Position Embedding Transformation for Multi-View 3D Object Detection

    Yingfei Liu, Tiancai Wang, Xiangyu Zhang +1

    cs.CVarXiv:2203.05625v32022
  55. Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution

    Haotian Tang, Zhijian Liu, Shengyu Zhao +4

    cs.CVarXiv:2007.16100v22020
  56. SpiderCNN: Deep Learning on Point Sets with Parameterized Convolutional Filters

    Yifan Xu, Tianqi Fan, Mingye Xu +2

    cs.CVarXiv:1803.11527v32018
  57. A Large-Scale Car Dataset for Fine-Grained Categorization and Verification

    Linjie Yang, Ping Luo, Chen Change Loy +1

    cs.CVcs.AIarXiv:1506.08959v22015
  58. PixelGen: Improving Pixel Diffusion with Perceptual Supervision

    Zehong Ma, Ruihan Xu, Shiliang Zhang

    cs.CVcs.AIarXiv:2602.02493v22026
  59. FILIP: Fine-grained Interactive Language-Image Pre-Training

    Lewei Yao, Runhui Huang, Lu Hou +7

    cs.CVcs.LGarXiv:2111.07783v12021
  60. SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models

    Jiwoo Chung, Sangeek Hyun, MinKyu Lee +5

    cs.CVarXiv:2602.18993v22026