Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,421 to 9,480 of 18,811

  1. Hand-Object Contact Consistency Reasoning for Human Grasps Generation

    Hanwen Jiang, Shaowei Liu, Jiashun Wang +1

    cs.CVarXiv:2104.03304v12021
  2. DO-Conv: Depthwise Over-parameterized Convolutional Layer

    Jinming Cao, Yangyan Li, Mingchao Sun +5

    cs.CVeess.IVarXiv:2006.12030v12020
  3. Neural Kinematic Networks for Unsupervised Motion Retargetting

    Ruben Villegas, Jimei Yang, Duygu Ceylan +1

    cs.CVarXiv:1804.05653v12018
  4. CogVLM2: Visual Language Models for Image and Video Understanding

    Wenyi Hong, Weihan Wang, Ming Ding +22

    cs.CVarXiv:2408.16500v12024
  5. Dual Contrastive Learning for General Face Forgery Detection

    Ke Sun, Taiping Yao, Shen Chen +3

    cs.CVarXiv:2112.13522v12021
  6. Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

    Daiqing Li, Aleks Kamko, Ehsan Akhgari +3

    cs.CVcs.AIarXiv:2402.17245v12024
  7. Forgetting Outside the Box: Scrubbing Deep Networks of Information Accessible from Input-Output Observations

    Aditya Golatkar, Alessandro Achille, Stefano Soatto

    cs.LGcs.CVcs.ITarXiv:2003.02960v32020
  8. Grasping Field: Learning Implicit Representations for Human Grasps

    Korrawe Karunratanakul, Jinlong Yang, Yan Zhang +3

    cs.CVarXiv:2008.04451v32020
  9. NaVILA: Legged Robot Vision-Language-Action Model for Navigation

    An-Chieh Cheng, Yandong Ji, Zhaojing Yang +7

    cs.ROcs.CVarXiv:2412.04453v22024
  10. Patch-VQ: 'Patching Up' the Video Quality Problem

    Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram +1

    cs.CVarXiv:2011.13544v22020
  11. InstanceCut: from Edges to Instances with MultiCut

    Alexander Kirillov, Evgeny Levinkov, Bjoern Andres +2

    cs.CVarXiv:1611.08272v12016
  12. 3D-MPA: Multi Proposal Aggregation for 3D Semantic Instance Segmentation

    Francis Engelmann, Martin Bokeloh, Alireza Fathi +2

    cs.CVarXiv:2003.13867v12020
  13. Stable and Controllable Neural Texture Synthesis and Style Transfer Using Histogram Losses

    Eric Risser, Pierre Wilmot, Connelly Barnes

    cs.GRcs.CVcs.NEarXiv:1701.08893v22017
  14. MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments

    Manolis Savva, Angel X. Chang, Alexey Dosovitskiy +2

    cs.LGcs.AIcs.CVarXiv:1712.03931v12017
  15. An Interpretable Deep Hierarchical Semantic Convolutional Neural Network for Lung Nodule Malignancy Classification

    Shiwen Shen, Simon X. Han, Denise R. Aberle +2

    cs.CVcs.AIarXiv:1806.00712v12018
  16. Top-push Video-based Person Re-identification

    Jinjie You, Ancong Wu, Xiang Li +1

    cs.CVarXiv:1604.08683v22016
  17. Interpretations are useful: penalizing explanations to align neural networks with prior knowledge

    Laura Rieger, Chandan Singh, W. James Murdoch +1

    cs.LGcs.CVstat.MLarXiv:1909.13584v42019
  18. A continual learning survey: Defying forgetting in classification tasks

    Matthias De Lange, Rahaf Aljundi, Marc Masana +5

    cs.CVstat.MLarXiv:1909.08383v32019
  19. Deep Recurrent Neural Network for Mobile Human Activity Recognition with High Throughput

    Masaya Inoue, Sozo Inoue, Takeshi Nishida

    cs.CVcs.NEarXiv:1611.03607v12016
  20. TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition

    Chih-Yao Ma, Min-Hung Chen, Zsolt Kira +1

    cs.CVarXiv:1703.10667v12017
  21. Binary Patterns Encoded Convolutional Neural Networks for Texture Recognition and Remote Sensing Scene Classification

    Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost van de Weijer +2

    cs.CVarXiv:1706.01171v22017
  22. Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models

    Rinon Gal, Moab Arar, Yuval Atzmon +3

    cs.CVcs.GRcs.LGarXiv:2302.12228v32023
  23. Joint Domain Alignment and Discriminative Feature Learning for Unsupervised Deep Domain Adaptation

    Chao Chen, Zhihong Chen, Boyuan Jiang +1

    cs.LGcs.CVstat.MLarXiv:1808.09347v22018
  24. 4D-Rotor Gaussian Splatting: Towards Efficient Novel View Synthesis for Dynamic Scenes

    Yuanxing Duan, Fangyin Wei, Qiyu Dai +3

    cs.CVarXiv:2402.03307v32024
  25. Discovery of Latent 3D Keypoints via End-to-end Geometric Reasoning

    Supasorn Suwajanakorn, Noah Snavely, Jonathan Tompson +1

    cs.CVcs.LGstat.MLarXiv:1807.03146v22018
  26. InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset

    Wenbin Li, Sajad Saeedi, John McCormac +6

    cs.CVcs.AIcs.LGarXiv:1809.00716v12018
  27. What Can Low Resource Languages Learn From Each Other?

    Achyuth P, Kahaan Shah, Chetan Arora

    cs.CVarXiv:2608.27753v12026
  28. Non-locally Enhanced Encoder-Decoder Network for Single Image De-raining

    Guanbin Li, Xiang He, Wei Zhang +3

    cs.CVarXiv:1808.01491v12018
  29. EfficientPS: Efficient Panoptic Segmentation

    Rohit Mohan, Abhinav Valada

    cs.CVcs.LGcs.ROarXiv:2004.02307v32020
  30. Transfer Adaptation Learning: A Decade Survey

    Lei Zhang, Xinbo Gao

    cs.CVarXiv:1903.04687v22019
  31. Attention Based Glaucoma Detection: A Large-scale Database and CNN Model

    Liu Li, Mai Xu, Xiaofei Wang +2

    cs.CVarXiv:1903.10831v32019
  32. RGB-T Image Saliency Detection via Collaborative Graph Learning

    Zhengzheng Tu, Tian Xia, Chenglong Li +3

    cs.CVarXiv:1905.06741v12019
  33. Generative Modeling using the Sliced Wasserstein Distance

    Ishan Deshpande, Ziyu Zhang, Alexander Schwing

    cs.CVarXiv:1803.11188v12018
  34. Uncertainty Modeling for Out-of-Distribution Generalization

    Xiaotong Li, Yongxing Dai, Yixiao Ge +3

    cs.CVcs.LGarXiv:2202.03958v22022
  35. mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

    Anwen Hu, Haiyang Xu, Jiabo Ye +8

    cs.CVarXiv:2403.12895v12024
  36. Shallow Triple Stream Three-dimensional CNN (STSTNet) for Micro-expression Recognition

    Sze-Teng Liong, Y. S. Gan, John See +2

    cs.CVarXiv:1902.03634v22019
  37. A Review of Object Detection Models based on Convolutional Neural Network

    F. Sultana, A. Sufian, P. Dutta

    cs.CVarXiv:1905.01614v32019
  38. Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting

    Yongqi Mao, Zijia Dai, Zhishuo Liu +3

    cs.CVarXiv:2608.28174v12026
  39. Fast and Robust Hand Tracking Using Detection-Guided Optimization

    Srinath Sridhar, Franziska Mueller, Antti Oulasvirta +1

    cs.CVarXiv:1602.04124v12016
  40. Recurrent Fully Convolutional Neural Networks for Multi-slice MRI Cardiac Segmentation

    Rudra P K Poudel, Pablo Lamata, Giovanni Montana

    stat.MLcs.CVcs.LGarXiv:1608.03974v12016
  41. Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models

    Xindi Yang, Yicheng Wu, Cheng Zhang +2

    cs.CVarXiv:2608.28082v12026
  42. Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics

    Ling-Yu Duan, Jiaying Liu, Wenhan Yang +2

    cs.CVarXiv:2001.03569v22020
  43. TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking

    Peng Chu, Jiang Wang, Quanzeng You +2

    cs.CVarXiv:2104.00194v22021
  44. FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba

    Xinyu Xie, Yawen Cui, Tao Tan +2

    cs.CVarXiv:2404.09498v32024
  45. SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization

    Yue Jiang, Dantong Ji, Zhizhong Han +1

    cs.CVcs.GRcs.LGarXiv:1912.07109v22019
  46. Explainable and Explicit Visual Reasoning over Scene Graphs

    Jiaxin Shi, Hanwang Zhang, Juanzi Li

    cs.CVarXiv:1812.01855v22018
  47. DeepCap: Monocular Human Performance Capture Using Weak Supervision

    Marc Habermann, Weipeng Xu, Michael Zollhoefer +2

    cs.CVarXiv:2003.08325v12020
  48. ABCNet: Attentive Bilateral Contextual Network for Efficient Semantic Segmentation of Fine-Resolution Remote Sensing Images

    Rui Li, Chenxi Duan

    cs.CVarXiv:2102.02531v12021
  49. Deep Learning for Face Anti-Spoofing: A Survey

    Zitong Yu, Yunxiao Qin, Xiaobai Li +3

    cs.CVarXiv:2106.14948v32021
  50. An Explainable Machine Learning Model for Early Detection of Parkinson's Disease using LIME on DaTscan Imagery

    Pavan Rajkumar Magesh, Richard Delwin Myloth, Rijo Jackson Tom

    cs.CVcs.LGeess.IVarXiv:2008.00238v12020
  51. Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images

    Xiaoman Zhang, Chaoyi Wu, Ya Zhang +2

    cs.CVarXiv:2302.14042v32023
  52. SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

    Tao Sun, Mattia Segu, Janis Postels +5

    cs.CVcs.LGarXiv:2206.08367v12022
  53. Learning Monocular 3D Human Pose Estimation from Multi-view Images

    Helge Rhodin, Jörg Spörri, Isinsu Katircioglu +5

    cs.CVarXiv:1803.04775v22018
  54. There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average

    Ben Athiwaratkun, Marc Finzi, Pavel Izmailov +1

    cs.LGcs.AIcs.CVarXiv:1806.05594v32018
  55. Face De-Spoofing: Anti-Spoofing via Noise Modeling

    Amin Jourabloo, Yaojie Liu, Xiaoming Liu

    cs.CVarXiv:1807.09968v12018
  56. The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes

    Abhishek Patil, Srikanth Malla, Haiming Gang +1

    cs.CVcs.ROarXiv:1903.01568v12019
  57. SNIPS: Solving Noisy Inverse Problems Stochastically

    Bahjat Kawar, Gregory Vaksman, Michael Elad

    eess.IVcs.CVarXiv:2105.14951v22021
  58. Understanding The Robustness in Vision Transformers

    Daquan Zhou, Zhiding Yu, Enze Xie +4

    cs.CVarXiv:2204.12451v42022
  59. Honeybee: Locality-enhanced Projector for Multimodal LLM

    Junbum Cha, Wooyoung Kang, Jonghwan Mun +1

    cs.CVcs.AIcs.CLarXiv:2312.06742v22023
  60. Extreme clicking for efficient object annotation

    Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller +1

    cs.CVarXiv:1708.02750v12017