Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,861 to 13,920 of 18,819

  1. Physical Adversarial Examples for Object Detectors

    Kevin Eykholt, Ivan Evtimov, Earlence Fernandes +6

    cs.CRcs.CVcs.LGarXiv:1807.07769v22018
  2. PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images

    Yingfei Liu, Junjie Yan, Fan Jia +5

    cs.CVarXiv:2206.01256v32022
  3. ABD-Net: Attentive but Diverse Person Re-Identification

    Tianlong Chen, Shaojin Ding, Jingyi Xie +5

    cs.CVarXiv:1908.01114v32019
  4. PAGS: Autofocusing Photoacoustic Tomography via Speed-of-Sound-Adaptive Gaussian Splatting

    Jiarui Ge, Jintao Ma, Bangxu Fan +4

    cs.CVphysics.med-pharXiv:2608.25472v12026
  5. VideoComposer: Compositional Video Synthesis with Motion Controllability

    Xiang Wang, Hangjie Yuan, Shiwei Zhang +6

    cs.CVarXiv:2306.02018v22023
  6. Learning to Learn Single Domain Generalization

    Fengchun Qiao, Long Zhao, Xi Peng

    cs.CVarXiv:2003.13216v12020
  7. RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

    Zhuoyan Liu, Yihan Wang, Bo Wang +2

    cs.CVarXiv:2608.25367v12026
  8. MakeItTalk: Speaker-Aware Talking-Head Animation

    Yang Zhou, Xintong Han, Eli Shechtman +3

    cs.CVcs.GRarXiv:2004.12992v32020
  9. YouTube-VOS: Sequence-to-Sequence Video Object Segmentation

    Ning Xu, Linjie Yang, Yuchen Fan +6

    cs.CVarXiv:1809.00461v12018
  10. Features for Multi-Target Multi-Camera Tracking and Re-Identification

    Ergys Ristani, Carlo Tomasi

    cs.CVarXiv:1803.10859v12018
  11. Wavelet Convolutions for Large Receptive Fields

    Shahaf E. Finder, Roy Amoyal, Eran Treister +1

    cs.CVarXiv:2407.05848v22024
  12. Ultra Fast Structure-aware Deep Lane Detection

    Zequn Qin, Huanyu Wang, Xi Li

    cs.CVarXiv:2004.11757v42020
  13. AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

    Xinze Liu, Lei Yang, Dayan Wu +7

    cs.CVarXiv:2608.25412v12026
  14. Learning to Generate Novel Domains for Domain Generalization

    Kaiyang Zhou, Yongxin Yang, Timothy Hospedales +1

    cs.CVarXiv:2007.03304v32020
  15. Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

    Jihun Kim, Hyun-Kurl Jang, Hyemin Yang +3

    cs.CVarXiv:2608.25418v12026
  16. Finite Scalar Quantization: VQ-VAE Made Simple

    Fabian Mentzer, David Minnen, Eirikur Agustsson +1

    cs.CVcs.LGarXiv:2309.15505v22023
  17. MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

    Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao +3

    cs.CVcs.AIarXiv:2608.25575v12026
  18. Large-Scale Adversarial Training for Vision-and-Language Representation Learning

    Zhe Gan, Yen-Chun Chen, Linjie Li +3

    cs.CVcs.CLcs.LGarXiv:2006.06195v22020
  19. PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

    Jingyang Su, Pu Cao, Xiuze Jin +3

    cs.CVarXiv:2608.25299v12026
  20. A survey of active learning algorithms for supervised remote sensing image classification

    Devis Tuia, Michele Volpi, Loris Copa +2

    cs.CVarXiv:2104.07784v12021
  21. MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

    Zihao Zhang, Dayan Wu, Xinze Liu +6

    cs.CVarXiv:2608.25305v12026
  22. Whole Slide Images based Cancer Survival Prediction using Attention Guided Deep Multiple Instance Learning Networks

    Jiawen Yao, Xinliang Zhu, Jitendra Jonnagaddala +2

    eess.IVcs.CVarXiv:2009.11169v12020
  23. MEMO: Test Time Robustness via Adaptation and Augmentation

    Marvin Zhang, Sergey Levine, Chelsea Finn

    cs.LGcs.CVarXiv:2110.09506v32021
  24. Learning Factorized Multimodal Representations

    Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh +2

    cs.LGcs.CLcs.CVarXiv:1806.06176v32018
  25. AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

    Xintong Zhang, Xiaomeng Fan, Shilin Yan +7

    cs.CVcs.AIarXiv:2608.25559v12026
  26. GraftSR: Grafting Authentic Textures for Real-World Image Super-Resolution via Identical-Instance Guidance

    Qifan Yu, Haoran Bai, Zongyao He +4

    cs.CVarXiv:2608.25334v12026
  27. Towards Stable Test-Time Adaptation in Dynamic Wild World

    Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang +4

    cs.LGcs.CVarXiv:2302.12400v12023
  28. MAMA-FLUX.2: Image-to-Image Synthesis of Post-Contrast Breast DCE-MRI for the MAMA-SYNTH Challenge

    Kamil Kwarciak, Marek Wodzinski

    cs.CVarXiv:2608.25648v12026
  29. HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis

    Xihui Liu, Haiyu Zhao, Maoqing Tian +5

    cs.CVarXiv:1709.09930v12017
  30. Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery

    Ali Lesani, Chul Min Yeum, Su-Min Kang

    cs.CVcs.AIcs.ROarXiv:2608.25435v12026
  31. Equalization Loss for Long-Tailed Object Recognition

    Jingru Tan, Changbao Wang, Buyu Li +4

    cs.CVarXiv:2003.05176v22020
  32. OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

    Zimin Xia, Mubariz Zaffar, Junsheng Fu +2

    cs.CVarXiv:2608.25274v12026
  33. Image to Image Translation for Domain Adaptation

    Zak Murez, Soheil Kolouri, David Kriegman +2

    cs.CVarXiv:1712.00479v12017
  34. Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

    Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke

    cs.CVcs.AIarXiv:2608.25495v12026
  35. AdaViT: Adaptive Tokens for Efficient Vision Transformer

    Hongxu Yin, Arash Vahdat, Jose Alvarez +3

    cs.CVcs.LGarXiv:2112.07658v32021
  36. Voxel Transformer for 3D Object Detection

    Jiageng Mao, Yujing Xue, Minzhe Niu +5

    cs.CVarXiv:2109.02497v22021
  37. CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos

    Kaiser Hamid, Can Cui, Nade Liang

    cs.CVcs.AIarXiv:2608.25344v12026
  38. Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

    Yiwen Liang, Hui Chen, Yizhe Xiong +7

    cs.CVarXiv:2608.25653v12026
  39. VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

    Longteng Jiang, DanDan Zheng, Qianqian Qiao +7

    cs.CVcs.AIarXiv:2608.25452v12026
  40. Token-Oriented Semantic Communication with Pretrained Vision Transformers

    Jiwoong Im, Minwoo Kim, Jaeho Lee +2

    eess.SPcs.AIcs.CVarXiv:2608.25410v12026
    Summaries:简体中文
  41. FateZero: Fusing Attentions for Zero-shot Text-based Video Editing

    Chenyang Qi, Xiaodong Cun, Yong Zhang +4

    cs.CVarXiv:2303.09535v32023
  42. Hierarchical MoE for Multi-Modal ILD Diagnosis

    Alec K. Peltekian, Gorkem Durak, Halil Ertugrul Aktas +9

    cs.AIcs.CVarXiv:2608.25261v12026
  43. A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

    Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov +14

    cs.CVcs.LGstat.MLarXiv:1910.04867v22019
  44. DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation

    Chuixuan Fan, Guang Li, Shijie Wang +5

    cs.CVcs.AIarXiv:2608.25480v12026
  45. U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

    Chenxin Li, Xinyu Liu, Wuyang Li +5

    eess.IVcs.CVarXiv:2406.02918v32024
  46. M3D-RPN: Monocular 3D Region Proposal Network for Object Detection

    Garrick Brazil, Xiaoming Liu

    cs.CVarXiv:1907.06038v22019
  47. TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-rays

    Xiaosong Wang, Yifan Peng, Le Lu +2

    cs.CVarXiv:1801.04334v12018
  48. OpenVeinNet: Robust Open-Set Finger Vein Verification with Dynamic Snake Convolution and Graph Learning

    Sushrut Patwardhan, Raghavendra Ramachandra

    cs.CVarXiv:2608.25515v12026
  49. Local Relation Networks for Image Recognition

    Han Hu, Zheng Zhang, Zhenda Xie +1

    cs.CVcs.AIcs.LGarXiv:1904.11491v12019
  50. SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving

    Tiago Cortinhal, George Tzelepis, Eren Erdal Aksoy

    cs.CVcs.LGarXiv:2003.03653v42020
  51. Semi-Supervised Adaptation of Vision-Language Models for Image Classification

    Mohamed L. Mekhalfi, Mohamad M. Al Rahhal, Yakoub Bazi +4

    cs.CVarXiv:2608.25485v12026
  52. MotionCLIP: Exposing Human Motion Generation to CLIP Space

    Guy Tevet, Brian Gordon, Amir Hertz +2

    cs.CVcs.GRarXiv:2203.08063v12022
  53. TD-MPC2: Scalable, Robust World Models for Continuous Control

    Nicklas Hansen, Hao Su, Xiaolong Wang

    cs.LGcs.AIcs.CVarXiv:2310.16828v22023
  54. HATS: Histograms of Averaged Time Surfaces for Robust Event-based Object Classification

    Amos Sironi, Manuele Brambilla, Nicolas Bourdis +2

    cs.CVarXiv:1803.07913v12018
  55. Convolutional Neural Networks Applied to House Numbers Digit Classification

    Pierre Sermanet, Soumith Chintala, Yann LeCun

    cs.CVcs.LGcs.NEarXiv:1204.3968v12012
  56. Large Separable Kernel Attention: Rethinking the Large Kernel Attention Design in CNN

    Kin Wai Lau, Lai-Man Po, Yasar Abbas Ur Rehman

    cs.CVarXiv:2309.01439v32023
  57. Normalized Loss Functions for Deep Learning with Noisy Labels

    Xingjun Ma, Hanxun Huang, Yisen Wang +3

    cs.LGcs.CVstat.MLarXiv:2006.13554v12020
  58. Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude Learning

    Yu Tian, Guansong Pang, Yuanhong Chen +3

    cs.CVarXiv:2101.10030v32021
  59. Privacy-preserving Federated Brain Tumour Segmentation

    Wenqi Li, Fausto Milletarì, Daguang Xu +8

    cs.CVarXiv:1910.00962v12019
  60. Suggestive Annotation: A Deep Active Learning Framework for Biomedical Image Segmentation

    Lin Yang, Yizhe Zhang, Jianxu Chen +2

    cs.CVarXiv:1706.04737v12017