Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,581 to 14,640 of 18,830

  1. ROI-Gated SAHI: Content-Adaptive Slicing-Based Inference for Efficient Object Detection

    Rashid Riyadh, Abd Ullah Khan, Imad Gohar +1

    cs.CVarXiv:2608.23923v12026
  2. Slimmable Neural Networks

    Jiahui Yu, Linjie Yang, Ning Xu +2

    cs.CVcs.AIarXiv:1812.08928v12018
  3. SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models

    Yuhang Su, Mei Wang, Yaoyao Zhong +4

    cs.CVcs.AIarXiv:2601.06944v12026
  4. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

    Liang Chen, Haozhe Zhao, Tianyu Liu +4

    cs.CVcs.AIcs.CLarXiv:2403.06764v32024
  5. GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models

    Zhankai Ye, Bofan Li, Yukai Jin +5

    cs.CVcs.AIarXiv:2601.07632v42026
  6. CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition

    Yuge Huang, Yuhan Wang, Ying Tai +5

    cs.CVarXiv:2004.00288v12020
  7. KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

    Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev +1

    cs.LGcs.AIcs.CVarXiv:2601.14232v22026
  8. Large Multimodal Models as General In-Context Classifiers

    Marco Garosi, Matteo Farina, Alessandro Conti +2

    cs.CVarXiv:2602.23229v12026
  9. Alterbute: Editing Intrinsic Attributes of Objects in Images

    Tal Reiss, Daniel Winter, Matan Cohen +4

    cs.CVcs.GRarXiv:2601.10714v22026
  10. CLEVRER: CoLlision Events for Video REpresentation and Reasoning

    Kexin Yi, Chuang Gan, Yunzhu Li +4

    cs.CVcs.AIcs.CLarXiv:1910.01442v22019
  11. Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identification

    Yixiao Ge, Dapeng Chen, Hongsheng Li

    cs.CVarXiv:2001.01526v22020
  12. Multi30K: Multilingual English-German Image Descriptions

    Desmond Elliott, Stella Frank, Khalil Sima'an +1

    cs.CLcs.CVarXiv:1605.00459v12016
  13. PredRNN: A Recurrent Neural Network for Spatiotemporal Predictive Learning

    Yunbo Wang, Haixu Wu, Jianjin Zhang +4

    cs.LGcs.CVarXiv:2103.09504v42021
  14. Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction

    Guy Gafni, Justus Thies, Michael Zollhöfer +1

    cs.CVcs.GRarXiv:2012.03065v12020
  15. Low-Rank Ternary Adaptation for Fine-Tuning Transformers

    Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert

    cs.CVcs.LGarXiv:2608.24469v12026
  16. CenterMask : Real-Time Anchor-Free Instance Segmentation

    Youngwan Lee, Jongyoul Park

    cs.CVarXiv:1911.06667v62019
  17. MOTS: Multi-Object Tracking and Segmentation

    Paul Voigtlaender, Michael Krause, Aljosa Osep +4

    cs.CVarXiv:1902.03604v22019
  18. CA-Net: Comprehensive Attention Convolutional Neural Networks for Explainable Medical Image Segmentation

    Ran Gu, Guotai Wang, Tao Song +6

    eess.IVcs.CVarXiv:2009.10549v22020
  19. Learning to Upsample by Learning to Sample

    Wenze Liu, Hao Lu, Hongtao Fu +1

    cs.CVarXiv:2308.15085v12023
  20. RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection

    Yilmaz Korkmaz, Vishal M. Patel

    cs.CVarXiv:2601.11898v12026
  21. Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison

    Dongxu Li, Cristian Rodriguez Opazo, Xin Yu +1

    cs.CVcs.HCcs.MMarXiv:1910.11006v22019
  22. DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

    Zhiyu Wu, Xiaokang Chen, Zizheng Pan +24

    cs.CVcs.AIcs.CLarXiv:2412.10302v12024
  23. Learning a Discriminative Null Space for Person Re-identification

    Li Zhang, Tao Xiang, Shaogang Gong

    cs.CVarXiv:1603.02139v12016
  24. AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients

    Juntang Zhuang, Tommy Tang, Yifan Ding +4

    cs.LGcs.CVstat.MLarXiv:2010.07468v52020
  25. Deep Unfolding Network for Image Super-Resolution

    Kai Zhang, Luc Van Gool, Radu Timofte

    eess.IVcs.CVarXiv:2003.10428v12020
  26. Rotational Projection Statistics for 3D Local Surface Description and Object Recognition

    Yulan Guo, Ferdous Sohel, Mohammed Bennamoun +2

    cs.CVarXiv:1304.3192v12013
  27. Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

    Tristan Thrush, Ryan Jiang, Max Bartolo +4

    cs.CVcs.CLarXiv:2204.03162v22022
  28. Siamese Instance Search for Tracking

    Ran Tao, Efstratios Gavves, Arnold W. M. Smeulders

    cs.CVarXiv:1605.05863v12016
  29. WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

    Rishi Upadhyay, Howard Zhang, Jim Solomon +5

    cs.CVarXiv:2601.21282v22026
  30. Physics-guided Neural Networks (PGNN): An Application in Lake Temperature Modeling

    Arka Daw, Anuj Karpatne, William Watkins +2

    cs.LGcs.AIcs.CVarXiv:1710.11431v32017
  31. SynthSeg: Segmentation of brain MRI scans of any contrast and resolution without retraining

    Benjamin Billot, Douglas N. Greve, Oula Puonti +5

    eess.IVcs.CVarXiv:2107.09559v42021
  32. Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

    Alperen Kantarci, Visvanathan Ramesh, Gemma Roig

    cs.CVcs.AIcs.HCarXiv:2608.24340v12026
  33. ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding

    Junyi Hu, Tian Bai, Fengyi Wu +3

    cs.CVarXiv:2601.22666v12026
  34. nocaps: novel object captioning at scale

    Harsh Agrawal, Karan Desai, Yufei Wang +7

    cs.CVcs.AIcs.CLarXiv:1812.08658v32018
  35. Zero-Shot Text-Guided Object Generation with Dream Fields

    Ajay Jain, Ben Mildenhall, Jonathan T. Barron +2

    cs.CVcs.AIcs.GRarXiv:2112.01455v22021
  36. Generating High-Quality Crowd Density Maps using Contextual Pyramid CNNs

    Vishwanath A. Sindagi, Vishal M. Patel

    cs.CVarXiv:1708.00953v12017
  37. ObjEmbed: Towards Universal Multimodal Object Embeddings

    Shenghao Fu, Yukun Su, Fengyun Rao +3

    cs.CVarXiv:2602.01753v32026
  38. Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning

    Yiming Sun, Bing Cao, Pengfei Zhu +1

    cs.CVcs.LGeess.IVarXiv:2003.02437v22020
  39. Siam R-CNN: Visual Tracking by Re-Detection

    Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr +1

    cs.CVarXiv:1911.12836v22019
  40. FOTBCD: A Large-Scale Building Change Detection Benchmark from French Orthophotos and Topographic Data

    Abdelrrahman Moubane

    cs.CVarXiv:2601.22596v12026
  41. Modular Primitives for High-Performance Differentiable Rendering

    Samuli Laine, Janne Hellsten, Tero Karras +3

    cs.GRcs.CVcs.LGarXiv:2011.03277v12020
  42. Semantic Image Segmentation via Deep Parsing Network

    Ziwei Liu, Xiaoxiao Li, Ping Luo +2

    cs.CVarXiv:1509.02634v22015
  43. Generative Adversarial Networks for Extreme Learned Image Compression

    Eirikur Agustsson, Michael Tschannen, Fabian Mentzer +2

    cs.CVcs.LGarXiv:1804.02958v32018
  44. Invertible Conditional GANs for image editing

    Guim Perarnau, Joost van de Weijer, Bogdan Raducanu +1

    cs.CVcs.AIarXiv:1611.06355v12016
  45. NativeTok: Native Visual Tokenization for Improved Image Generation

    Bin Wu, Mengqi Huang, Weinan Jia +1

    cs.CVarXiv:2601.22837v12026
  46. The Deepfake Detection Challenge (DFDC) Preview Dataset

    Brian Dolhansky, Russ Howes, Ben Pflaum +2

    cs.CVcs.CYarXiv:1910.08854v22019
  47. IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

    Feyza Yavuz, Mert Bülent Sarıyıldız, Diane Larlus

    cs.CVarXiv:2608.24759v12026
  48. Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General

    Christoffer Koo Øhrstrøm, Rafael I. Cabral Muchacho, Yifei Dong +4

    cs.CVcs.LGarXiv:2602.01418v22026
  49. Luce: Relightable Gaussians for 3D Asset Generation

    Mayank Singh, Michele Stoppa, Alvise Memo +7

    cs.CVcs.AIcs.GRarXiv:2608.23943v12026
  50. BCN20000: Dermoscopic Lesions in the Wild

    Marc Combalia, Noel C. F. Codella, Veronica Rotemberg +8

    eess.IVcs.CVarXiv:1908.02288v22019
  51. Implicit neural representation of textures

    Albert Kwok, Zheyuan Hu, Dounia Hammou

    cs.CVcs.AIcs.GRarXiv:2602.02354v12026
  52. Knockoff Nets: Stealing Functionality of Black-Box Models

    Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

    cs.CVcs.CRcs.LGarXiv:1812.02766v12018
  53. EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models

    Yu Bai, MingMing Yu, Chaojie Li +3

    cs.ROcs.CVarXiv:2602.04515v12026
  54. Segment Anything in High Quality

    Lei Ke, Mingqiao Ye, Martin Danelljan +4

    cs.CVarXiv:2306.01567v22023
  55. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    Bo Miao, Weijia Liu, Jun Luo +8

    cs.CVcs.ROarXiv:2602.02220v22026
  56. DETRs with Collaborative Hybrid Assignments Training

    Zhuofan Zong, Guanglu Song, Yu Liu

    cs.CVarXiv:2211.12860v62022
  57. XCiT: Cross-Covariance Image Transformers

    Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron +8

    cs.CVcs.LGarXiv:2106.09681v22021
  58. Masked Autoencoders As Spatiotemporal Learners

    Christoph Feichtenhofer, Haoqi Fan, Yanghao Li +1

    cs.CVcs.LGarXiv:2205.09113v22022
  59. DeiT III: Revenge of the ViT

    Hugo Touvron, Matthieu Cord, Hervé Jégou

    cs.CVarXiv:2204.07118v12022
  60. Source-Face Authenticity Detection for 3D Gaussian Heads Reconstructed from a Single Portrait: A Benchmark and Dedicated Detector

    Yujie Gao, Zijian Yu, Yan Hong +2

    cs.CVarXiv:2608.23984v12026