Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,161 to 14,220 of 18,830

  1. SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

    Long Shu, Shuochen Liu, Wei Chen +4

    cs.CVcs.AIarXiv:2608.21796v12026
  2. Deep Subspace Clustering Networks

    Pan Ji, Tong Zhang, Hongdong Li +2

    cs.CVarXiv:1709.02508v12017
  3. GhostNetV2: Enhance Cheap Operation with Long-Range Attention

    Yehui Tang, Kai Han, Jianyuan Guo +3

    cs.CVarXiv:2211.12905v12022
  4. Latte: Latent Diffusion Transformer for Video Generation

    Xin Ma, Yaohui Wang, Xinyuan Chen +5

    cs.CVarXiv:2401.03048v32024
  5. Mode Regularized Generative Adversarial Networks

    Tong Che, Yanran Li, Athul Paul Jacob +2

    cs.LGcs.AIcs.CVarXiv:1612.02136v52016
  6. ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting

    Zongsheng Yue, Jianyi Wang, Chen Change Loy

    cs.CVarXiv:2307.12348v32023
  7. The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

    Liangzhi Li, Bowen Wang, Yiming Qian +3

    cs.CVcs.LGarXiv:2608.23634v12026
  8. LF-Net: Learning Local Features from Images

    Yuki Ono, Eduard Trulls, Pascal Fua +1

    cs.CVarXiv:1805.09662v22018
  9. Unsupervised Discovery of Mid-Level Discriminative Patches

    Saurabh Singh, Abhinav Gupta, Alexei A. Efros

    cs.CVcs.AIcs.LGarXiv:1205.3137v22012
  10. clDice -- A Novel Topology-Preserving Loss Function for Tubular Structure Segmentation

    Suprosanna Shit, Johannes C. Paetzold, Anjany Sekuboyina +6

    cs.CVcs.LGeess.IVarXiv:2003.07311v72020
  11. OCGAN: One-class Novelty Detection Using GANs with Constrained Latent Representations

    Pramuditha Perera, Ramesh Nallapati, Bing Xiang

    cs.CVcs.LGarXiv:1903.08550v12019
  12. SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge

    JeongRae Kim, Chaehyun Kim, Changwon Lim

    cs.CVarXiv:2608.22193v12026
  13. Image2StyleGAN++: How to Edit the Embedded Images?

    Rameen Abdal, Yipeng Qin, Peter Wonka

    cs.CVcs.GRarXiv:1911.11544v22019
  14. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng +5

    cs.CVarXiv:2403.14627v22024
  15. You Only Learn One Representation: Unified Network for Multiple Tasks

    Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao

    cs.CVarXiv:2105.04206v12021
  16. Ternary Weight Networks

    Fengfu Li, Bin Liu, Xiaoxing Wang +2

    cs.CVarXiv:1605.04711v32016
  17. MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation

    Xucong Zhang, Yusuke Sugano, Mario Fritz +1

    cs.CVarXiv:1711.09017v12017
  18. BC-IHV: Conditioning the Color Space for Stable Rectified-Flow Low-Light Enhancement

    Yi Ai, Zheng Chen, Yuanhao Cai +2

    cs.CVarXiv:2608.21847v12026
  19. Robust fine-tuning of zero-shot models

    Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim +8

    cs.CVcs.LGarXiv:2109.01903v32021
    Summaries:한국어
  20. Real-Time Adaptive Image Compression

    Oren Rippel, Lubomir Bourdev

    stat.MLcs.CVcs.LGarXiv:1705.05823v12017
  21. Progressively Learning Heterogeneous Skills in a Unified Latent Space

    Yue-Yi Zhang, Ming Gong, Linpu He +2

    cs.CVcs.AIcs.ROarXiv:2608.23258v12026
  22. HeatTok: Enhancing Remote Sensing Image Understanding via Thermodiffusion-based Tokenization

    Yingying Yan, Jiaqi Tang, Wei Wei +6

    cs.CVarXiv:2608.22485v12026
  23. Abnormal Event Detection in Videos using Spatiotemporal Autoencoder

    Yong Shean Chong, Yong Haur Tay

    cs.CVarXiv:1701.01546v12017
  24. SAL: Sign Agnostic Learning of Shapes from Raw Data

    Matan Atzmon, Yaron Lipman

    cs.CVcs.GRcs.LGarXiv:1911.10414v22019
  25. Cycle-Dehaze: Enhanced CycleGAN for Single Image Dehazing

    Deniz Engin, Anıl Genç, Hazım Kemal Ekenel

    cs.CVarXiv:1805.05308v12018
  26. Revisiting Local Descriptor based Image-to-Class Measure for Few-shot Learning

    Wenbin Li, Lei Wang, Jinglin Xu +3

    cs.CVarXiv:1903.12290v22019
  27. Curriculum Learning: A Survey

    Petru Soviany, Radu Tudor Ionescu, Paolo Rota +1

    cs.LGcs.CLcs.CVarXiv:2101.10382v32021
  28. Gradient Harmonized Single-stage Detector

    Buyu Li, Yu Liu, Xiaogang Wang

    cs.CVarXiv:1811.05181v12018
  29. Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation

    Pan Zhang, Bo Zhang, Ting Zhang +3

    cs.CVarXiv:2101.10979v22021
  30. ByteAction: Byte-space Action Recognition Foundation Model

    Fangcheng Li, Zhen Yu, Kejun Wu +2

    cs.CVarXiv:2608.22760v12026
  31. DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation

    Yujie Qi, Luyan Zhang

    cs.CVarXiv:2608.22885v12026
  32. T-LESS: An RGB-D Dataset for 6D Pose Estimation of Texture-less Objects

    Tomas Hodan, Pavel Haluza, Stepan Obdrzalek +3

    cs.CVcs.AIcs.ROarXiv:1701.05498v12017
  33. RGB-T Object Tracking:Benchmark and Baseline

    Chenglong Li, Xinyan Liang, Yijuan Lu +2

    cs.CVarXiv:1805.08982v12018
  34. Deep Learning for Classification of Hyperspectral Data: A Comparative Review

    Nicolas Audebert, Bertrand Saux, Sébastien Lefèvre

    cs.LGcs.CVcs.NEarXiv:1904.10674v12019
  35. Deep Spatial Autoencoders for Visuomotor Learning

    Chelsea Finn, Xin Yu Tan, Yan Duan +3

    cs.LGcs.CVcs.ROarXiv:1509.06113v32015
  36. Deep Neural Networks Improve Radiologists' Performance in Breast Cancer Screening

    Nan Wu, Jason Phang, Jungkyu Park +29

    cs.LGcs.CVstat.MLarXiv:1903.08297v12019
  37. Deep Learning Ensembles for Melanoma Recognition in Dermoscopy Images

    Noel Codella, Quoc-Bao Nguyen, Sharath Pankanti +4

    cs.CVarXiv:1610.04662v22016
  38. Bayesian CP Factorization of Incomplete Tensors with Automatic Rank Determination

    Qibin Zhao, Liqing Zhang, Andrzej Cichocki

    cs.LGcs.CVstat.MLarXiv:1401.6497v22014
  39. Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet

    Wieland Brendel, Matthias Bethge

    cs.CVcs.LGstat.MLarXiv:1904.00760v12019
  40. The Sound of Pixels

    Hang Zhao, Chuang Gan, Andrew Rouditchenko +3

    cs.CVcs.SDeess.ASarXiv:1804.03160v42018
  41. PixelLink: Detecting Scene Text via Instance Segmentation

    Dan Deng, Haifeng Liu, Xuelong Li +1

    cs.CVarXiv:1801.01315v12018
  42. Score-based diffusion models for accelerated MRI

    Hyungjin Chung, Jong Chul Ye

    eess.IVcs.AIcs.CVarXiv:2110.05243v32021
  43. SegAN: Adversarial Network with Multi-scale $L_1$ Loss for Medical Image Segmentation

    Yuan Xue, Tao Xu, Han Zhang +2

    cs.CVarXiv:1706.01805v22017
  44. Dual-Path Convolutional Image-Text Embeddings with Instance Loss

    Zhedong Zheng, Liang Zheng, Michael Garrett +3

    cs.CVcs.MMarXiv:1711.05535v42017
  45. Modelling Uncertainty in Deep Learning for Camera Relocalization

    Alex Kendall, Roberto Cipolla

    cs.CVcs.ROarXiv:1509.05909v22015
  46. PersonLab: Person Pose Estimation and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model

    George Papandreou, Tyler Zhu, Liang-Chieh Chen +3

    cs.CVarXiv:1803.08225v12018
  47. More Motion Is Not Always Better Motion: Corpus Composition Governs Whether Augmentation Helps SMPL-Based Parkinsonian Gait Severity Estimation

    Michael Caiola, Andrew C. Weitz

    cs.CVarXiv:2608.23730v12026
  48. M$^3$ISR: A Multi-Modal Multi-View Benchmark for 3D/4D Gaussian Splatting and Feedforward Compression

    Xinhui Liu, Lei Liu, Zhenghao Chen +3

    cs.CVarXiv:2608.22465v12026
  49. Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video

    Jia-Wang Bian, Zhichao Li, Naiyan Wang +4

    cs.CVarXiv:1908.10553v22019
  50. Hyper^2: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency

    Guantian Zheng, Haiyang Xu, Tianyu Gao

    cs.CVarXiv:2608.22238v12026
  51. DeepMVS: Learning Multi-view Stereopsis

    Po-Han Huang, Kevin Matzen, Johannes Kopf +2

    cs.CVarXiv:1804.00650v12018
  52. Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

    Renrui Zhang, Rongyao Fang, Wei Zhang +5

    cs.CVcs.CLarXiv:2111.03930v22021
  53. Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation

    Jordi Pont-Tuset, Pablo Arbelaez, Jonathan T. Barron +2

    cs.CVarXiv:1503.00848v42015
  54. Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs

    Haithem Turki, Deva Ramanan, Mahadev Satyanarayanan

    cs.CVcs.GRcs.LGarXiv:2112.10703v22021
  55. Visual Translation Embedding Network for Visual Relation Detection

    Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang +1

    cs.CVarXiv:1702.08319v12017
  56. Visual Reinforcement Learning with Imagined Goals

    Ashvin Nair, Vitchyr Pong, Murtaza Dalal +3

    cs.LGcs.CVcs.ROarXiv:1807.04742v22018
  57. Restoring Without Forgetting: Continual Learning Across Image Degradations

    Alif Ashrafee, Bartosz Krawczyk

    cs.CVcs.AIcs.LGarXiv:2608.23799v12026
  58. BehaviorWorldGen: Closing the Loop between Action Models and World Simulators via Controllable Behavior-Aware Structured World Generation

    Jiaqi Wang, Zhuo Zhang, Haining Guan +15

    cs.ROcs.CVarXiv:2608.22187v12026
  59. Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions

    Ayaan Haque, Matthew Tancik, Alexei A. Efros +2

    cs.CVcs.GRarXiv:2303.12789v22023
  60. Transferable Joint Attribute-Identity Deep Learning for Unsupervised Person Re-Identification

    Jingya Wang, Xiatian Zhu, Shaogang Gong +1

    cs.CVarXiv:1803.09786v12018