Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,761 to 11,820 of 18,817

  1. FedVision: An Online Visual Object Detection Platform Powered by Federated Learning

    Yang Liu, Anbu Huang, Yun Luo +7

    cs.LGcs.CVstat.MLarXiv:2001.06202v12020
  2. SELF: Learning to Filter Noisy Labels with Self-Ensembling

    Duc Tam Nguyen, Chaithanya Kumar Mummadi, Thi Phuong Nhung Ngo +3

    cs.CVcs.LGstat.MLarXiv:1910.01842v12019
  3. Deep Learning Methods for Parallel Magnetic Resonance Image Reconstruction

    Florian Knoll, Kerstin Hammernik, Chi Zhang +4

    eess.SPcs.CVcs.LGarXiv:1904.01112v12019
  4. InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

    Pan Zhang, Xiaoyi Dong, Bin Wang +18

    cs.CVarXiv:2309.15112v52023
  5. Learning to Reconstruct People in Clothing from a Single RGB Camera

    Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar +2

    cs.CVarXiv:1903.05885v22019
  6. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

    Hao Shao, Shengju Qian, Han Xiao +5

    cs.CVarXiv:2403.16999v32024
  7. CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval

    Zihao Wang, Xihui Liu, Hongsheng Li +4

    cs.CVarXiv:1909.05506v12019
  8. Transformers are Sample-Efficient World Models

    Vincent Micheli, Eloi Alonso, François Fleuret

    cs.LGcs.AIcs.CVarXiv:2209.00588v22022
  9. Constructing Unrestricted Adversarial Examples with Generative Models

    Yang Song, Rui Shu, Nate Kushman +1

    cs.LGcs.AIcs.CRarXiv:1805.07894v42018
  10. Learning to Diversify for Single Domain Generalization

    Zijian Wang, Yadan Luo, Ruihong Qiu +2

    cs.CVarXiv:2108.11726v32021
  11. Effective Face Frontalization in Unconstrained Images

    Tal Hassner, Shai Harel, Eran Paz +1

    cs.CVarXiv:1411.7964v12014
  12. Learning to Regress 3D Face Shape and Expression from an Image without 3D Supervision

    Soubhik Sanyal, Timo Bolkart, Haiwen Feng +1

    cs.CVarXiv:1905.06817v12019
  13. Cooperative Perception for 3D Object Detection in Driving Scenarios using Infrastructure Sensors

    Eduardo Arnold, Mehrdad Dianati, Robert de Temple +1

    cs.CVcs.LGcs.MAarXiv:1912.12147v22019
  14. Very high resolution canopy height maps from RGB imagery using self-supervised vision transformer and convolutional decoder trained on Aerial Lidar

    Jamie Tolan, Hung-I Yang, Ben Nosarzewski +13

    cs.CVcs.LGarXiv:2304.07213v32023
  15. Embracing Single Stride 3D Object Detector with Sparse Transformer

    Lue Fan, Ziqi Pang, Tianyuan Zhang +5

    cs.CVarXiv:2112.06375v12021
  16. Prompt Distribution Learning

    Yuning Lu, Jianzhuang Liu, Yonggang Zhang +2

    cs.CVarXiv:2205.03340v12022
  17. One Network to Solve Them All --- Solving Linear Inverse Problems using Deep Projection Models

    J. H. Rick Chang, Chun-Liang Li, Barnabas Poczos +2

    cs.CVarXiv:1703.09912v12017
  18. Learning Aberrance Repressed Correlation Filters for Real-Time UAV Tracking

    Ziyuan Huang, Changhong Fu, Yiming Li +2

    cs.CVarXiv:1908.02231v22019
  19. Iterative Learning with Open-set Noisy Labels

    Yisen Wang, Weiyang Liu, Xingjun Ma +4

    cs.CVarXiv:1804.00092v12018
  20. Whole-Body Human Pose Estimation in the Wild

    Sheng Jin, Lumin Xu, Jin Xu +5

    cs.CVarXiv:2007.11858v12020
  21. TempCompass: Do Video LLMs Really Understand Videos?

    Yuanxin Liu, Shicheng Li, Yi Liu +6

    cs.CVarXiv:2403.00476v32024
  22. No More Discrimination: Cross City Adaptation of Road Scene Segmenters

    Yi-Hsin Chen, Wei-Yu Chen, Yu-Ting Chen +3

    cs.CVcs.AIarXiv:1704.08509v12017
  23. Multi-adversarial Faster-RCNN for Unrestricted Object Detection

    Zhenwei He, Lei Zhang

    cs.CVarXiv:1907.10343v22019
  24. Local-Global Video-Text Interactions for Temporal Grounding

    Jonghwan Mun, Minsu Cho, Bohyung Han

    cs.CVarXiv:2004.07514v12020
  25. Progressive Pose Attention Transfer for Person Image Generation

    Zhen Zhu, Tengteng Huang, Baoguang Shi +3

    cs.CVarXiv:1904.03349v32019
  26. The Devil Is in the Details: Window-based Attention for Image Compression

    Renjie Zou, Chunfeng Song, Zhaoxiang Zhang

    eess.IVcs.CVarXiv:2203.08450v12022
  27. 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation

    Angela Dai, Matthias Nießner

    cs.CVarXiv:1803.10409v12018
  28. BAGAN: Data Augmentation with Balancing GAN

    Giovanni Mariani, Florian Scheidegger, Roxana Istrate +2

    cs.CVcs.LGstat.MLarXiv:1803.09655v22018
  29. Efficient Video Object Segmentation via Network Modulation

    Linjie Yang, Yanran Wang, Xuehan Xiong +2

    cs.CVarXiv:1802.01218v12018
  30. DC-SPP-YOLO: Dense Connection and Spatial Pyramid Pooling Based YOLO for Object Detection

    Zhanchao Huang, Jianlin Wang, Xuesong Fu +3

    cs.CVarXiv:1903.08589v22019
  31. Multi-Directional Multi-Level Dual-Cross Patterns for Robust Face Recognition

    Changxing Ding, Jonghyun Choi, Dacheng Tao +1

    cs.CVarXiv:1401.5311v22014
  32. Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry

    Nan Yang, Rui Wang, Jörg Stückler +1

    cs.CVarXiv:1807.02570v22018
  33. i-RevNet: Deep Invertible Networks

    Jörn-Henrik Jacobsen, Arnold Smeulders, Edouard Oyallon

    cs.LGcs.CVstat.MLarXiv:1802.07088v12018
  34. A Survey on 3D Gaussian Splatting

    Guikun Chen, Wenguan Wang

    cs.CVcs.AIcs.GRarXiv:2401.03890v92024
  35. BBDM: Image-to-image Translation with Brownian Bridge Diffusion Models

    Bo Li, Kaitao Xue, Bin Liu +1

    cs.CVeess.IVarXiv:2205.07680v22022
  36. Hand Pose Estimation via Latent 2.5D Heatmap Regression

    Umar Iqbal, Pavlo Molchanov, Thomas Breuel +2

    cs.CVcs.LGarXiv:1804.09534v12018
  37. DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

    Bencheng Liao, Shaoyu Chen, Haoran Yin +8

    cs.CVcs.ROarXiv:2411.15139v32024
  38. Deep Projective 3D Semantic Segmentation

    Felix Järemo Lawin, Martin Danelljan, Patrik Tosteberg +3

    cs.CVarXiv:1705.03428v12017
  39. EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain

    Wei Zhang, Miaoxin Cai, Tong Zhang +2

    cs.CVarXiv:2401.16822v32024
  40. Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis

    Seunghoon Hong, Dingdong Yang, Jongwook Choi +1

    cs.CVarXiv:1801.05091v22018
  41. Real-World Robot Learning with Masked Visual Pre-training

    Ilija Radosavovic, Tete Xiao, Stephen James +3

    cs.ROcs.CVcs.LGarXiv:2210.03109v12022
  42. Point2Sequence: Learning the Shape Representation of 3D Point Clouds with an Attention-based Sequence to Sequence Network

    Xinhai Liu, Zhizhong Han, Yu-Shen Liu +1

    cs.CVarXiv:1811.02565v22018
  43. Clinically Accurate Chest X-Ray Report Generation

    Guanxiong Liu, Tzu-Ming Harry Hsu, Matthew McDermott +4

    cs.CVcs.CLarXiv:1904.02633v22019
  44. The Pose Knows: Video Forecasting by Generating Pose Futures

    Jacob Walker, Kenneth Marino, Abhinav Gupta +1

    cs.CVarXiv:1705.00053v12017
  45. GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting

    Kai Zhang, Sai Bi, Hao Tan +4

    cs.CVarXiv:2404.19702v12024
  46. Understanding and Improving Fast Adversarial Training

    Maksym Andriushchenko, Nicolas Flammarion

    cs.LGcs.CRcs.CVarXiv:2007.02617v22020
  47. Anomaly Detection in Video via Self-Supervised and Multi-Task Learning

    Mariana-Iuliana Georgescu, Antonio Barbalau, Radu Tudor Ionescu +3

    cs.CVcs.LGeess.IVarXiv:2011.07491v32020
  48. Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition

    Tianshui Chen, Muxin Xu, Xiaolu Hui +2

    cs.CVarXiv:1908.07325v12019
  49. Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation

    Shiqi Yang, Yaxing Wang, Joost van de Weijer +2

    cs.CVcs.LGarXiv:2110.04202v32021
  50. From BoW to CNN: Two Decades of Texture Representation for Texture Classification

    Li Liu, Jie Chen, Paul Fieguth +3

    cs.CVcs.LGarXiv:1801.10324v22018
  51. Real-World Single Image Super-Resolution: A Brief Review

    Honggang Chen, Xiaohai He, Linbo Qing +3

    eess.IVcs.CVarXiv:2103.02368v12021
  52. Low-Light Image Enhancement with Wavelet-based Diffusion Models

    Hai Jiang, Ao Luo, Songchen Han +2

    cs.CVarXiv:2306.00306v32023
  53. Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

    Yiyang Zhou, Chenhang Cui, Jaehong Yoon +5

    cs.LGcs.CLcs.CVarXiv:2310.00754v22023
  54. Video Object Segmentation Without Temporal Information

    Kevis-Kokitsi Maninis, Sergi Caelles, Yuhua Chen +4

    cs.CVarXiv:1709.06031v22017
  55. One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion

    Minghua Liu, Ruoxi Shi, Linghao Chen +7

    cs.CVcs.AIcs.GRarXiv:2311.07885v12023
  56. OpenMask3D: Open-Vocabulary 3D Instance Segmentation

    Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner +3

    cs.CVarXiv:2306.13631v22023
  57. Gaussian Temporal Awareness Networks for Action Localization

    Fuchen Long, Ting Yao, Zhaofan Qiu +3

    cs.CVarXiv:1909.03877v12019
  58. Source Data-absent Unsupervised Domain Adaptation through Hypothesis Transfer and Labeling Transfer

    Jian Liang, Dapeng Hu, Yunbo Wang +2

    cs.CVcs.LGarXiv:2012.07297v32020
  59. A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction

    Zhian Liu, Yongwei Nie, Chengjiang Long +2

    cs.CVarXiv:2108.06852v12021
  60. Large Scale Image Completion via Co-Modulated Generative Adversarial Networks

    Shengyu Zhao, Jonathan Cui, Yilun Sheng +4

    cs.CVcs.GRcs.LGarXiv:2103.10428v12021