Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,121 to 12,180 of 18,830

  1. Designing Network Design Strategies Through Gradient Path Analysis

    Chien-Yao Wang, Hong-Yuan Mark Liao, I-Hau Yeh

    cs.CVarXiv:2211.04800v12022
  2. Motion Representations for Articulated Animation

    Aliaksandr Siarohin, Oliver J. Woodford, Jian Ren +2

    cs.CVarXiv:2104.11280v12021
  3. Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks

    Yu-Gang Jiang, Zuxuan Wu, Jun Wang +2

    cs.CVcs.MMarXiv:1502.07209v22015
  4. Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network

    Xinjing Cheng, Peng Wang, Ruigang Yang

    cs.CVarXiv:1808.00150v12018
  5. Lipschitz-Margin Training: Scalable Certification of Perturbation Invariance for Deep Neural Networks

    Yusuke Tsuzuku, Issei Sato, Masashi Sugiyama

    cs.CVcs.LGstat.MLarXiv:1802.04034v32018
  6. Visual Instance Retrieval with Deep Convolutional Networks

    Ali Sharif Razavian, Josephine Sullivan, Stefan Carlsson +1

    cs.CVarXiv:1412.6574v42014
  7. SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

    An-Chieh Cheng, Hongxu Yin, Yang Fu +5

    cs.CVarXiv:2406.01584v32024
  8. Instance-level Human Parsing via Part Grouping Network

    Ke Gong, Xiaodan Liang, Yicheng Li +3

    cs.CVarXiv:1808.00157v12018
  9. Weakly Supervised Action Localization by Sparse Temporal Pooling Network

    Phuc Nguyen, Ting Liu, Gautam Prasad +1

    cs.CVarXiv:1712.05080v22017
  10. What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector

    Muhammad Yaseen

    cs.CVarXiv:2408.15857v12024
  11. Addressing Failure Prediction by Learning Model Confidence

    Charles Corbière, Nicolas Thome, Avner Bar-Hen +2

    cs.CVcs.LGstat.MLarXiv:1910.04851v22019
  12. ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas +4

    cs.CVarXiv:2204.13662v32022
  13. Real-Time Scene Text Detection with Differentiable Binarization and Adaptive Scale Fusion

    Minghui Liao, Zhisheng Zou, Zhaoyi Wan +2

    cs.CVarXiv:2202.10304v12022
  14. PointDSC: Robust Point Cloud Registration using Deep Spatial Consistency

    Xuyang Bai, Zixin Luo, Lei Zhou +5

    cs.CVarXiv:2103.05465v12021
  15. Geo-LoRA: Geometry-Aware Subspace Evolution for Low-Rank Adaptation in Continual Learning

    Yibo Feng

    cs.CVarXiv:2608.26960v12026
  16. From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

    Tianwei Yin, Qiang Zhang, Richard Zhang +4

    cs.CVarXiv:2412.07772v42024
  17. Application of Deep Convolutional Neural Networks for Detecting Extreme Weather in Climate Datasets

    Yunjie Liu, Evan Racah, Prabhat +6

    cs.CVarXiv:1605.01156v12016
  18. Deep Stereo using Adaptive Thin Volume Representation with Uncertainty Awareness

    Shuo Cheng, Zexiang Xu, Shilin Zhu +4

    cs.CVcs.LGcs.ROarXiv:1911.12012v22019
  19. Hull First, Wake Second: Wake-Reliance Suppression for Robust Maritime Vessel Detection

    Yefan Wang, Xingyu Wang, Ruibiao Zhu +1

    cs.CVarXiv:2608.26665v12026
  20. NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

    Daniel Dauner, Marcel Hallgarten, Tianyu Li +9

    cs.CVcs.AIcs.LGarXiv:2406.15349v22024
  21. Deep convolutional neural networks for brain image analysis on magnetic resonance imaging: a review

    Jose Bernal, Kaisar Kushibar, Daniel S. Asfaw +4

    cs.CVarXiv:1712.03747v32017
  22. 3D-VLA: A 3D Vision-Language-Action Generative World Model

    Haoyu Zhen, Xiaowen Qiu, Peihao Chen +5

    cs.CVcs.AIcs.CLarXiv:2403.09631v12024
  23. SmartBrush: Text and Shape Guided Object Inpainting with Diffusion Model

    Shaoan Xie, Zhifei Zhang, Zhe Lin +2

    cs.CVarXiv:2212.05034v12022
  24. Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

    Haihan Li, Haihao Li, Zhenfei Xu +1

    cs.CVarXiv:2608.26866v12026
  25. Multi-View Intact Space Learning

    Chang Xu, Dacheng Tao, Chao Xu

    cs.CVarXiv:1904.02340v12019
  26. Rethinking Image Processing for the Age of AI: A Problem-First Framework for Scientific Progress

    Guoping Qiu

    cs.CVarXiv:2608.26833v12026
  27. UniSim: A Neural Closed-Loop Sensor Simulator

    Ze Yang, Yun Chen, Jingkang Wang +4

    cs.CVcs.ROarXiv:2308.01898v12023
  28. DeepSD: Generating High Resolution Climate Change Projections through Single Image Super-Resolution

    Thomas Vandal, Evan Kodra, Sangram Ganguly +3

    cs.CVarXiv:1703.03126v12017
  29. Evaluator-Dependent Patient-Adaptive ECG Lead-Channel Allocation

    Xiaoyang Li, Zeyan Tao

    cs.CVarXiv:2608.26827v12026
  30. Language2Pose: Natural Language Grounded Pose Forecasting

    Chaitanya Ahuja, Louis-Philippe Morency

    cs.CVcs.CLarXiv:1907.01108v22019
  31. I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

    Shiwei Zhang, Jiayu Wang, Yingya Zhang +6

    cs.CVarXiv:2311.04145v12023
  32. Attribute Prototype Network for Zero-Shot Learning

    Wenjia Xu, Yongqin Xian, Jiuniu Wang +2

    cs.CVcs.LGarXiv:2008.08290v42020
  33. PieAPP: Perceptual Image-Error Assessment through Pairwise Preference

    Ekta Prashnani, Hong Cai, Yasamin Mostofi +1

    cs.CVarXiv:1806.02067v12018
  34. Generative Semantic Scene Completion

    Shi Chen, Weifeng Ge

    cs.CVcs.LGcs.ROarXiv:2608.26737v12026
  35. Domain-Specific Self-Supervised Representation Learning for Retinal Fundus Classification

    Bekzat Nurlanbekova, Fung Fung Ting

    cs.CVcs.LGarXiv:2608.26686v12026
  36. pix2code: Generating Code from a Graphical User Interface Screenshot

    Tony Beltramelli

    cs.LGcs.AIcs.CLarXiv:1705.07962v22017
  37. What matters when building vision-language models?

    Hugo Laurençon, Léo Tronchon, Matthieu Cord +1

    cs.CVcs.AIarXiv:2405.02246v12024
  38. SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

    Haizhao Fan, Xinyi Le

    cs.LGcs.CVarXiv:2608.26829v12026
  39. Sketch-based 3D Shape Retrieval using Convolutional Neural Networks

    Fang Wang, Le Kang, Yi Li

    cs.CVarXiv:1504.03504v12015
  40. UIEC^2-Net: CNN-based Underwater Image Enhancement Using Two Color Space

    Yudong Wang, Jichang Guo, Huan Gao +1

    cs.CVarXiv:2103.07138v22021
  41. CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

    Wenxiao Wang, Lu Yao, Long Chen +4

    cs.CVcs.LGarXiv:2108.00154v22021
  42. ARCH: Animatable Reconstruction of Clothed Humans

    Zeng Huang, Yuanlu Xu, Christoph Lassner +2

    cs.GRcs.CVcs.LGarXiv:2004.04572v22020
  43. Multiple Futures Prediction

    Yichuan Charlie Tang, Ruslan Salakhutdinov

    cs.LGcs.CVcs.MAarXiv:1911.00997v22019
  44. Spatially Adaptive Computation Time for Residual Networks

    Michael Figurnov, Maxwell D. Collins, Yukun Zhu +4

    cs.CVcs.LGarXiv:1612.02297v22016
  45. Aligning Domain-specific Distribution and Classifier for Cross-domain Classification from Multiple Sources

    Yongchun Zhu, Fuzhen Zhuang, Deqing Wang

    cs.LGcs.AIcs.CVarXiv:2201.01003v12022
  46. ClusterAttention: A training-free speedup of bidirectional attention

    Kasper Nordenram, Amelie Dittmann

    cs.LGcs.CVarXiv:2608.26965v12026
  47. Exploring Visual Prompts for Adapting Large-Scale Models

    Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan +1

    cs.CVarXiv:2203.17274v22022
  48. S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental Learning

    Yabin Wang, Zhiwu Huang, Xiaopeng Hong

    cs.CVcs.LGarXiv:2207.12819v22022
  49. Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

    Yuanduo Hong, Huihui Pan, Weichao Sun +1

    cs.CVarXiv:2101.06085v22021
  50. Hands Deep in Deep Learning for Hand Pose Estimation

    Markus Oberweger, Paul Wohlhart, Vincent Lepetit

    cs.CVarXiv:1502.06807v22015
  51. Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models

    Jiuxiang Gu, Jianfei Cai, Shafiq Joty +2

    cs.CVarXiv:1711.06420v22017
  52. Deep Direct Regression for Multi-Oriented Scene Text Detection

    Wenhao He, Xu-Yao Zhang, Fei Yin +1

    cs.CVarXiv:1703.08289v12017
  53. Implicit Diffusion Models for Continuous Super-Resolution

    Sicheng Gao, Xuhui Liu, Bohan Zeng +6

    cs.CVarXiv:2303.16491v22023
  54. Exploring the Landscape of Spatial Robustness

    Logan Engstrom, Brandon Tran, Dimitris Tsipras +2

    cs.LGcs.CVcs.NEarXiv:1712.02779v42017
  55. LEDNet: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation

    Yu Wang, Quan Zhou, Jia Liu +4

    cs.CVarXiv:1905.02423v32019
  56. GridMask Data Augmentation

    Pengguang Chen, Shu Liu, Hengshuang Zhao +2

    cs.CVarXiv:2001.04086v32020
  57. Delta-encoder: an effective sample synthesis method for few-shot object recognition

    Eli Schwartz, Leonid Karlinsky, Joseph Shtok +6

    cs.CVarXiv:1806.04734v32018
  58. Learning the Model Update for Siamese Trackers

    Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer +2

    cs.CVarXiv:1908.00855v22019
  59. 3DFeat-Net: Weakly Supervised Local 3D Features for Point Cloud Registration

    Zi Jian Yew, Gim Hee Lee

    cs.CVarXiv:1807.09413v12018
  60. Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs

    Ruigang Fu, Qingyong Hu, Xiaohu Dong +3

    cs.CVcs.AIcs.LGarXiv:2008.02312v42020