Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,801 to 13,860 of 18,776

  1. Features for Multi-Target Multi-Camera Tracking and Re-Identification

    Ergys Ristani, Carlo Tomasi

    cs.CVarXiv:1803.10859v12018
  2. Wavelet Convolutions for Large Receptive Fields

    Shahaf E. Finder, Roy Amoyal, Eran Treister +1

    cs.CVarXiv:2407.05848v22024
  3. Ultra Fast Structure-aware Deep Lane Detection

    Zequn Qin, Huanyu Wang, Xi Li

    cs.CVarXiv:2004.11757v42020
  4. AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval

    Xinze Liu, Lei Yang, Dayan Wu +7

    cs.CVarXiv:2608.25412v12026
  5. Learning to Generate Novel Domains for Domain Generalization

    Kaiyang Zhou, Yongxin Yang, Timothy Hospedales +1

    cs.CVarXiv:2007.03304v32020
  6. Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

    Jihun Kim, Hyun-Kurl Jang, Hyemin Yang +3

    cs.CVarXiv:2608.25418v12026
  7. Finite Scalar Quantization: VQ-VAE Made Simple

    Fabian Mentzer, David Minnen, Eirikur Agustsson +1

    cs.CVcs.LGarXiv:2309.15505v22023
  8. MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

    Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao +3

    cs.CVcs.AIarXiv:2608.25575v12026
  9. Large-Scale Adversarial Training for Vision-and-Language Representation Learning

    Zhe Gan, Yen-Chun Chen, Linjie Li +3

    cs.CVcs.CLcs.LGarXiv:2006.06195v22020
  10. PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

    Jingyang Su, Pu Cao, Xiuze Jin +3

    cs.CVarXiv:2608.25299v12026
  11. A survey of active learning algorithms for supervised remote sensing image classification

    Devis Tuia, Michele Volpi, Loris Copa +2

    cs.CVarXiv:2104.07784v12021
  12. MulVec: Fine-Grained Role-Aware Matching for Training-Free Zero-Shot Composed Image Retrieval

    Zihao Zhang, Dayan Wu, Xinze Liu +6

    cs.CVarXiv:2608.25305v12026
  13. Whole Slide Images based Cancer Survival Prediction using Attention Guided Deep Multiple Instance Learning Networks

    Jiawen Yao, Xinliang Zhu, Jitendra Jonnagaddala +2

    eess.IVcs.CVarXiv:2009.11169v12020
  14. MEMO: Test Time Robustness via Adaptation and Augmentation

    Marvin Zhang, Sergey Levine, Chelsea Finn

    cs.LGcs.CVarXiv:2110.09506v32021
  15. Learning Factorized Multimodal Representations

    Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh +2

    cs.LGcs.CLcs.CVarXiv:1806.06176v32018
  16. AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

    Xintong Zhang, Xiaomeng Fan, Shilin Yan +7

    cs.CVcs.AIarXiv:2608.25559v12026
  17. GraftSR: Grafting Authentic Textures for Real-World Image Super-Resolution via Identical-Instance Guidance

    Qifan Yu, Haoran Bai, Zongyao He +4

    cs.CVarXiv:2608.25334v12026
  18. Towards Stable Test-Time Adaptation in Dynamic Wild World

    Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang +4

    cs.LGcs.CVarXiv:2302.12400v12023
  19. MAMA-FLUX.2: Image-to-Image Synthesis of Post-Contrast Breast DCE-MRI for the MAMA-SYNTH Challenge

    Kamil Kwarciak, Marek Wodzinski

    cs.CVarXiv:2608.25648v12026
  20. HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis

    Xihui Liu, Haiyu Zhao, Maoqing Tian +5

    cs.CVarXiv:1709.09930v12017
  21. Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery

    Ali Lesani, Chul Min Yeum, Su-Min Kang

    cs.CVcs.AIcs.ROarXiv:2608.25435v12026
  22. Equalization Loss for Long-Tailed Object Recognition

    Jingru Tan, Changbao Wang, Buyu Li +4

    cs.CVarXiv:2003.05176v22020
  23. OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

    Zimin Xia, Mubariz Zaffar, Junsheng Fu +2

    cs.CVarXiv:2608.25274v12026
  24. Image to Image Translation for Domain Adaptation

    Zak Murez, Soheil Kolouri, David Kriegman +2

    cs.CVarXiv:1712.00479v12017
  25. Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

    Lewis de Zoete Grundy, Chris McCarthy, Christopher Fluke

    cs.CVcs.AIarXiv:2608.25495v12026
  26. AdaViT: Adaptive Tokens for Efficient Vision Transformer

    Hongxu Yin, Arash Vahdat, Jose Alvarez +3

    cs.CVcs.LGarXiv:2112.07658v32021
  27. Voxel Transformer for 3D Object Detection

    Jiageng Mao, Yujing Xue, Minzhe Niu +5

    cs.CVarXiv:2109.02497v22021
  28. CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos

    Kaiser Hamid, Can Cui, Nade Liang

    cs.CVcs.AIarXiv:2608.25344v12026
  29. Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

    Yiwen Liang, Hui Chen, Yizhe Xiong +7

    cs.CVarXiv:2608.25653v12026
  30. VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

    Longteng Jiang, DanDan Zheng, Qianqian Qiao +7

    cs.CVcs.AIarXiv:2608.25452v12026
  31. Token-Oriented Semantic Communication with Pretrained Vision Transformers

    Jiwoong Im, Minwoo Kim, Jaeho Lee +2

    eess.SPcs.AIcs.CVarXiv:2608.25410v12026
    Summaries:简体中文
  32. FateZero: Fusing Attentions for Zero-shot Text-based Video Editing

    Chenyang Qi, Xiaodong Cun, Yong Zhang +4

    cs.CVarXiv:2303.09535v32023
  33. Hierarchical MoE for Multi-Modal ILD Diagnosis

    Alec K. Peltekian, Gorkem Durak, Halil Ertugrul Aktas +9

    cs.AIcs.CVarXiv:2608.25261v12026
  34. A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

    Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov +14

    cs.CVcs.LGstat.MLarXiv:1910.04867v22019
  35. DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation

    Chuixuan Fan, Guang Li, Shijie Wang +5

    cs.CVcs.AIarXiv:2608.25480v12026
  36. U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

    Chenxin Li, Xinyu Liu, Wuyang Li +5

    eess.IVcs.CVarXiv:2406.02918v32024
  37. M3D-RPN: Monocular 3D Region Proposal Network for Object Detection

    Garrick Brazil, Xiaoming Liu

    cs.CVarXiv:1907.06038v22019
  38. TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-rays

    Xiaosong Wang, Yifan Peng, Le Lu +2

    cs.CVarXiv:1801.04334v12018
  39. OpenVeinNet: Robust Open-Set Finger Vein Verification with Dynamic Snake Convolution and Graph Learning

    Sushrut Patwardhan, Raghavendra Ramachandra

    cs.CVarXiv:2608.25515v12026
  40. Local Relation Networks for Image Recognition

    Han Hu, Zheng Zhang, Zhenda Xie +1

    cs.CVcs.AIcs.LGarXiv:1904.11491v12019
  41. SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving

    Tiago Cortinhal, George Tzelepis, Eren Erdal Aksoy

    cs.CVcs.LGarXiv:2003.03653v42020
  42. Semi-Supervised Adaptation of Vision-Language Models for Image Classification

    Mohamed L. Mekhalfi, Mohamad M. Al Rahhal, Yakoub Bazi +4

    cs.CVarXiv:2608.25485v12026
  43. MotionCLIP: Exposing Human Motion Generation to CLIP Space

    Guy Tevet, Brian Gordon, Amir Hertz +2

    cs.CVcs.GRarXiv:2203.08063v12022
  44. TD-MPC2: Scalable, Robust World Models for Continuous Control

    Nicklas Hansen, Hao Su, Xiaolong Wang

    cs.LGcs.AIcs.CVarXiv:2310.16828v22023
  45. HATS: Histograms of Averaged Time Surfaces for Robust Event-based Object Classification

    Amos Sironi, Manuele Brambilla, Nicolas Bourdis +2

    cs.CVarXiv:1803.07913v12018
  46. Convolutional Neural Networks Applied to House Numbers Digit Classification

    Pierre Sermanet, Soumith Chintala, Yann LeCun

    cs.CVcs.LGcs.NEarXiv:1204.3968v12012
  47. Large Separable Kernel Attention: Rethinking the Large Kernel Attention Design in CNN

    Kin Wai Lau, Lai-Man Po, Yasar Abbas Ur Rehman

    cs.CVarXiv:2309.01439v32023
  48. Normalized Loss Functions for Deep Learning with Noisy Labels

    Xingjun Ma, Hanxun Huang, Yisen Wang +3

    cs.LGcs.CVstat.MLarXiv:2006.13554v12020
  49. Weakly-supervised Video Anomaly Detection with Robust Temporal Feature Magnitude Learning

    Yu Tian, Guansong Pang, Yuanhong Chen +3

    cs.CVarXiv:2101.10030v32021
  50. Privacy-preserving Federated Brain Tumour Segmentation

    Wenqi Li, Fausto Milletarì, Daguang Xu +8

    cs.CVarXiv:1910.00962v12019
  51. Suggestive Annotation: A Deep Active Learning Framework for Biomedical Image Segmentation

    Lin Yang, Yizhe Zhang, Jianxu Chen +2

    cs.CVarXiv:1706.04737v12017
  52. Few-Shot Class-Incremental Learning

    Xiaoyu Tao, Xiaopeng Hong, Xinyuan Chang +3

    cs.CVcs.LGstat.MLarXiv:2004.10956v22020
  53. Maintaining Discrimination and Fairness in Class Incremental Learning

    Bowen Zhao, Xi Xiao, Guojun Gan +2

    cs.CVarXiv:1911.07053v12019
  54. GANerated Hands for Real-time 3D Hand Tracking from Monocular RGB

    Franziska Mueller, Florian Bernard, Oleksandr Sotnychenko +4

    cs.CVarXiv:1712.01057v12017
  55. HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

    Linjie Li, Yen-Chun Chen, Yu Cheng +3

    cs.CVcs.CLcs.LGarXiv:2005.00200v22020
  56. TEA: Temporal Excitation and Aggregation for Action Recognition

    Yan Li, Bin Ji, Xintian Shi +3

    cs.CVarXiv:2004.01398v12020
  57. Fast Segment Anything

    Xu Zhao, Wenchao Ding, Yongqi An +5

    cs.CVcs.AIarXiv:2306.12156v12023
  58. Unsupervised Post-Training of Foundation Models: A Survey

    Yijie Xu, Qianyi Cai, Huizai Yao +9

    cs.CLcs.AIcs.CVarXiv:2608.24982v12026
  59. A Patient-Centric Dataset of Images and Metadata for Identifying Melanomas Using Clinical Context

    Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein +21

    eess.IVcs.CVcs.CYarXiv:2008.07360v12020
  60. Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

    Hanbyul Joo, Tomas Simon, Yaser Sheikh

    cs.CVarXiv:1801.01615v12018