Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,681 to 13,740 of 18,916

  1. GeoChat: Grounded Large Vision-Language Model for Remote Sensing

    Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer +3

    cs.CVcs.AIarXiv:2311.15826v12023
  2. How Much Can CLIP Benefit Vision-and-Language Tasks?

    Sheng Shen, Liunian Harold Li, Hao Tan +5

    cs.CVcs.AIcs.CLarXiv:2107.06383v12021
  3. GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting

    Chi Yan, Delin Qu, Dan Xu +4

    cs.CVarXiv:2311.11700v42023
  4. Women also Snowboard: Overcoming Bias in Captioning Models

    Kaylee Burns, Lisa Anne Hendricks, Kate Saenko +2

    cs.CVarXiv:1803.09797v42018
  5. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets

    Longwen Zhang, Ziyu Wang, Qixuan Zhang +6

    cs.CVarXiv:2406.13897v12024
  6. Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates

    Leslie N. Smith, Nicholay Topin

    cs.LGcs.CVcs.NEarXiv:1708.07120v32017
  7. Emergent Correspondence from Image Diffusion

    Luming Tang, Menglin Jia, Qianqian Wang +2

    cs.CVarXiv:2306.03881v22023
  8. HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation

    Jiefeng Li, Chao Xu, Zhicun Chen +3

    cs.CVarXiv:2011.14672v42020
  9. Prompt-aligned Gradient for Prompt Tuning

    Beier Zhu, Yulei Niu, Yucheng Han +2

    cs.CVarXiv:2205.14865v42022
  10. T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos

    Kai Kang, Hongsheng Li, Junjie Yan +8

    cs.CVarXiv:1604.02532v42016
  11. Beyond Bilinear: Generalized Multimodal Factorized High-order Pooling for Visual Question Answering

    Zhou Yu, Jun Yu, Chenchao Xiang +2

    cs.CVarXiv:1708.03619v22017
  12. Learning Feature Pyramids for Human Pose Estimation

    Wei Yang, Shuang Li, Wanli Ouyang +2

    cs.CVarXiv:1708.01101v12017
    Summaries:한국어
  13. ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning

    Shengchao Hu, Li Chen, Penghao Wu +3

    cs.CVarXiv:2207.07601v22022
  14. Meta-SR: A Magnification-Arbitrary Network for Super-Resolution

    Xuecai Hu, Haoyuan Mu, Xiangyu Zhang +3

    cs.CVarXiv:1903.00875v42019
  15. Multi-Level Factorisation Net for Person Re-Identification

    Xiaobin Chang, Timothy M. Hospedales, Tao Xiang

    cs.CVarXiv:1803.09132v22018
  16. Detecting Visual Relationships with Deep Relational Networks

    Bo Dai, Yuqi Zhang, Dahua Lin

    cs.CVarXiv:1704.03114v22017
  17. Patch SVDD: Patch-level SVDD for Anomaly Detection and Segmentation

    Jihun Yi, Sungroh Yoon

    cs.CVarXiv:2006.16067v22020
  18. Grounding of Textual Phrases in Images by Reconstruction

    Anna Rohrbach, Marcus Rohrbach, Ronghang Hu +2

    cs.CVcs.CLcs.LGarXiv:1511.03745v42015
  19. An Uncertain Future: Forecasting from Static Images using Variational Autoencoders

    Jacob Walker, Carl Doersch, Abhinav Gupta +1

    cs.CVarXiv:1606.07873v12016
  20. A probabilistic atlas of the human thalamic nuclei combining ex vivo MRI and histology

    Juan Eugenio Iglesias, Ricardo Insausti, Garikoitz Lerma-Usabiaga +7

    q-bio.NCcs.CVphysics.med-pharXiv:1806.08634v12018
  21. WaveletKernelNet: An Interpretable Deep Neural Network for Industrial Intelligent Diagnosis

    Tianfu Li, Zhibin Zhao, Chuang Sun +4

    cs.CVcs.LGcs.NEarXiv:1911.07925v32019
  22. Fast inference of deep neural networks in FPGAs for particle physics

    Javier Duarte, Song Han, Philip Harris +8

    physics.ins-detcs.CVhep-exarXiv:1804.06913v32018
  23. ActionCLIP: A New Paradigm for Video Action Recognition

    Mengmeng Wang, Jiazheng Xing, Yong Liu

    cs.CVarXiv:2109.08472v12021
  24. Multimodal Motion Prediction with Stacked Transformers

    Yicheng Liu, Jinghuai Zhang, Liangji Fang +2

    cs.CVcs.AIarXiv:2103.11624v22021
  25. Star-convex Polyhedra for 3D Object Detection and Segmentation in Microscopy

    Martin Weigert, Uwe Schmidt, Robert Haase +2

    cs.CVarXiv:1908.03636v22019
  26. Evaluating Text-to-Visual Generation with Image-to-Text Generation

    Zhiqiu Lin, Deepak Pathak, Baiqi Li +5

    cs.CVcs.AIcs.CLarXiv:2404.01291v22024
  27. RankIQA: Learning from Rankings for No-reference Image Quality Assessment

    Xialei Liu, Joost van de Weijer, Andrew D. Bagdanov

    cs.CVarXiv:1707.08347v12017
  28. FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

    Guillaume Jaume, Hazim Kemal Ekenel, Jean-Philippe Thiran

    cs.IRcs.CVcs.LGarXiv:1905.13538v22019
  29. Rewrite the Stars

    Xu Ma, Xiyang Dai, Yue Bai +2

    cs.CVarXiv:2403.19967v12024
  30. Neural Body Fitting: Unifying Deep Learning and Model-Based Human Pose and Shape Estimation

    Mohamed Omran, Christoph Lassner, Gerard Pons-Moll +2

    cs.CVarXiv:1808.05942v12018
  31. Long-Term Feature Banks for Detailed Video Understanding

    Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan +3

    cs.CVarXiv:1812.05038v22018
  32. Edge-labeling Graph Neural Network for Few-shot Learning

    Jongmin Kim, Taesup Kim, Sungwoong Kim +1

    cs.LGcs.CVarXiv:1905.01436v12019
  33. The Curse of Recursion: Training on Generated Data Makes Models Forget

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao +3

    cs.LGcs.AIcs.CLarXiv:2305.17493v32023
  34. Articulated Pose Estimation by a Graphical Model with Image Dependent Pairwise Relations

    Xianjie Chen, Alan Yuille

    cs.CVarXiv:1407.3399v22014
  35. Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal Effect

    Kaihua Tang, Jianqiang Huang, Hanwang Zhang

    cs.CVcs.LGstat.MLarXiv:2009.12991v52020
  36. Unsupervised Medical Image Translation with Adversarial Diffusion Models

    Muzaffer Özbey, Onat Dalmaz, Salman UH Dar +4

    eess.IVcs.CVarXiv:2207.08208v32022
  37. Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units

    Wenling Shang, Kihyuk Sohn, Diogo Almeida +1

    cs.LGcs.CVarXiv:1603.05201v22016
  38. Learning to Find Good Correspondences

    Kwang Moo Yi, Eduard Trulls, Yuki Ono +3

    cs.CVarXiv:1711.05971v22017
  39. Tensor Robust Principal Component Analysis: Exact Recovery of Corrupted Low-Rank Tensors via Convex Optimization

    Canyi Lu, Jiashi Feng, Yudong Chen +3

    cs.CVarXiv:1708.04181v32017
  40. Fractional Max-Pooling

    Benjamin Graham

    cs.CVarXiv:1412.6071v42014
  41. A Survey on Instance Segmentation: State of the art

    Abdul Mueed Hafiz, Ghulam Mohiuddin Bhat

    cs.CVcs.LGeess.IVarXiv:2007.00047v12020
  42. F-Cooper: Feature based Cooperative Perception for Autonomous Vehicle Edge Computing System Using 3D Point Clouds

    Qi Chen

    cs.CVarXiv:1909.06459v12019
  43. Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models

    Ozan Özdenizci, Robert Legenstein

    cs.CVcs.LGarXiv:2207.14626v22022
  44. Cross-Age LFW: A Database for Studying Cross-Age Face Recognition in Unconstrained Environments

    Tianyue Zheng, Weihong Deng, Jiani Hu

    cs.CVcs.DBarXiv:1708.08197v12017
  45. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

    Xiang Yue, Tianyu Zheng, Yuansheng Ni +10

    cs.CLcs.CVarXiv:2409.02813v32024
  46. NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and Localization

    Qi Wang, Junyu Gao, Wei Lin +1

    cs.CVarXiv:2001.03360v42020
  47. ECO: Efficient Convolutional Network for Online Video Understanding

    Mohammadreza Zolfaghari, Kamaljeet Singh, Thomas Brox

    cs.CVcs.AIcs.IRarXiv:1804.09066v22018
  48. DesnowNet: Context-Aware Deep Network for Snow Removal

    Yun-Fu Liu, Da-Wei Jaw, Shih-Chia Huang +1

    cs.CVarXiv:1708.04512v12017
  49. 3D Morphable Face Models -- Past, Present and Future

    Bernhard Egger, William A. P. Smith, Ayush Tewari +10

    cs.CVcs.GRcs.LGarXiv:1909.01815v22019
  50. Scribbler: Controlling Deep Image Synthesis with Sketch and Color

    Patsorn Sangkloy, Jingwan Lu, Chen Fang +2

    cs.CVcs.LGarXiv:1612.00835v22016
  51. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

    Yi Wang, Yinan He, Yizhuo Li +13

    cs.CVarXiv:2307.06942v22023
  52. DepGraph: Towards Any Structural Pruning

    Gongfan Fang, Xinyin Ma, Mingli Song +2

    cs.AIcs.CVarXiv:2301.12900v22023
  53. Generative Multimodal Models are In-Context Learners

    Quan Sun, Yufeng Cui, Xiaosong Zhang +8

    cs.CVarXiv:2312.13286v22023
  54. Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases

    Christopher Clark, Mark Yatskar, Luke Zettlemoyer

    cs.CLcs.CVcs.LGarXiv:1909.03683v12019
  55. PoseTrack: A Benchmark for Human Pose Estimation and Tracking

    Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov +4

    cs.CVarXiv:1710.10000v22017
  56. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

    Jiale Xu, Weihao Cheng, Yiming Gao +3

    cs.CVarXiv:2404.07191v22024
  57. DeeperGCN: All You Need to Train Deeper GCNs

    Guohao Li, Chenxin Xiong, Ali Thabet +1

    cs.LGcs.CVstat.MLarXiv:2006.07739v12020
  58. LangSplat: 3D Language Gaussian Splatting

    Minghan Qin, Wanhua Li, Jiawei Zhou +2

    cs.CVarXiv:2312.16084v22023
  59. PARE: Part Attention Regressor for 3D Human Body Estimation

    Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges +1

    cs.CVarXiv:2104.08527v22021
  60. Pixel Difference Networks for Efficient Edge Detection

    Zhuo Su, Wenzhe Liu, Zitong Yu +5

    cs.CVarXiv:2108.07009v12021