Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

11,701 to 11,760 of 18,817

  1. A Deeper Look at Dataset Bias

    Tatiana Tommasi, Novi Patricia, Barbara Caputo +1

    cs.CVarXiv:1505.01257v12015
  2. LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

    Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao +14

    cs.CVarXiv:2410.17434v12024
  3. SoftGroup for 3D Instance Segmentation on Point Clouds

    Thang Vu, Kookhoi Kim, Tung M. Luu +2

    cs.CVarXiv:2203.01509v12022
  4. RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer

    Wenyu Lv, Yian Zhao, Qinyao Chang +3

    cs.CVarXiv:2407.17140v12024
  5. Regularizing Class-wise Predictions via Self-knowledge Distillation

    Sukmin Yun, Jongjin Park, Kimin Lee +1

    cs.LGcs.CVstat.MLarXiv:2003.13964v22020
  6. Semi-supervised Left Atrium Segmentation with Mutual Consistency Training

    Yicheng Wu, Minfeng Xu, Zongyuan Ge +2

    cs.CVarXiv:2103.02911v22021
  7. Just Ask: Learning to Answer Questions from Millions of Narrated Videos

    Antoine Yang, Antoine Miech, Josef Sivic +2

    cs.CVcs.CLcs.LGarXiv:2012.00451v32020
  8. Combining Residual Networks with LSTMs for Lipreading

    Themos Stafylakis, Georgios Tzimiropoulos

    cs.CVarXiv:1703.04105v42017
  9. SinSR: Diffusion-Based Image Super-Resolution in a Single Step

    Yufei Wang, Wenhan Yang, Xinyuan Chen +7

    cs.CVarXiv:2311.14760v12023
  10. RepMet: Representative-based metric learning for classification and one-shot object detection

    Leonid Karlinsky, Joseph Shtok, Sivan Harary +5

    cs.CVarXiv:1806.04728v32018
  11. Multi-Content GAN for Few-Shot Font Style Transfer

    Samaneh Azadi, Matthew Fisher, Vladimir Kim +3

    cs.CVarXiv:1712.00516v12017
  12. Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy

    Asit Mishra, Debbie Marr

    cs.LGcs.CVcs.NEarXiv:1711.05852v12017
  13. Evaluation of Algorithms for Multi-Modality Whole Heart Segmentation: An Open-Access Grand Challenge

    Xiahai Zhuang, Lei Li, Christian Payer +31

    cs.CVarXiv:1902.07880v12019
  14. Cross-modality Person re-identification with Shared-Specific Feature Transfer

    Yan Lu, Yue Wu, Bin Liu +4

    cs.CVarXiv:2002.12489v32020
  15. Reconstruction Network for Video Captioning

    Bairui Wang, Lin Ma, Wei Zhang +1

    cs.CVarXiv:1803.11438v12018
  16. Action Tubelet Detector for Spatio-Temporal Action Localization

    Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari +1

    cs.CVarXiv:1705.01861v32017
  17. Domain-Symmetric Networks for Adversarial Domain Adaptation

    Yabin Zhang, Hui Tang, Kui Jia +1

    cs.CVarXiv:1904.04663v22019
  18. Virtual to Real Reinforcement Learning for Autonomous Driving

    Xinlei Pan, Yurong You, Ziyan Wang +1

    cs.AIcs.CVarXiv:1704.03952v42017
  19. Keeping the Bad Guys Out: Protecting and Vaccinating Deep Learning with JPEG Compression

    Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen +4

    cs.CVcs.CRarXiv:1705.02900v12017
  20. NeX: Real-time View Synthesis with Neural Basis Expansion

    Suttisak Wizadwongsa, Pakkapon Phongthawee, Jiraphon Yenphraphai +1

    cs.CVcs.GRcs.LGarXiv:2103.05606v22021
  21. Shadow Detection: A Survey and Comparative Evaluation of Recent Methods

    Andres Sanin, Conrad Sanderson, Brian C. Lovell

    cs.CVcs.ROarXiv:1304.1233v12013
  22. When2com: Multi-Agent Perception via Communication Graph Grouping

    Yen-Cheng Liu, Junjiao Tian, Nathaniel Glaser +1

    cs.CVcs.MAcs.ROarXiv:2006.00176v22020
  23. Guided Image Generation with Conditional Invertible Neural Networks

    Lynton Ardizzone, Carsten Lüth, Jakob Kruse +2

    cs.CVcs.LGarXiv:1907.02392v32019
  24. Post-training Quantization on Diffusion Models

    Yuzhang Shang, Zhihang Yuan, Bin Xie +2

    cs.CVarXiv:2211.15736v32022
  25. Dynamic Channel Pruning: Feature Boosting and Suppression

    Xitong Gao, Yiren Zhao, Łukasz Dudziak +2

    cs.CVarXiv:1810.05331v22018
  26. Atlas: End-to-End 3D Scene Reconstruction from Posed Images

    Zak Murez, Tarrence van As, James Bartolozzi +3

    cs.CVarXiv:2003.10432v32020
  27. AutoFormer: Searching Transformers for Visual Recognition

    Minghao Chen, Houwen Peng, Jianlong Fu +1

    cs.CVarXiv:2107.00651v12021
  28. Seeing What a GAN Cannot Generate

    David Bau, Jun-Yan Zhu, Jonas Wulff +4

    cs.CVcs.GRcs.LGarXiv:1910.11626v12019
  29. Pushing the Boundaries of View Extrapolation with Multiplane Images

    Pratul P. Srinivasan, Richard Tucker, Jonathan T. Barron +3

    cs.CVarXiv:1905.00413v12019
  30. Track Anything: Segment Anything Meets Videos

    Jinyu Yang, Mingqi Gao, Zhe Li +3

    cs.CVarXiv:2304.11968v22023
  31. Video Compression through Image Interpolation

    Chao-Yuan Wu, Nayan Singhal, Philipp Krähenbühl

    cs.CVarXiv:1804.06919v12018
  32. Evaluating Scalable Bayesian Deep Learning Methods for Robust Computer Vision

    Fredrik K. Gustafsson, Martin Danelljan, Thomas B. Schön

    cs.LGcs.CVstat.MLarXiv:1906.01620v32019
  33. Crowd Counting and Density Estimation by Trellis Encoder-Decoder Network

    Xiaolong Jiang, Zehao Xiao, Baochang Zhang +4

    cs.CVarXiv:1903.00853v22019
  34. A comprehensive survey on point cloud registration

    Xiaoshui Huang, Guofeng Mei, Jian Zhang +1

    cs.CVarXiv:2103.02690v22021
  35. End-to-end Lane Shape Prediction with Transformers

    Ruijin Liu, Zejian Yuan, Tie Liu +1

    cs.CVcs.AIarXiv:2011.04233v22020
  36. SelfReg: Self-supervised Contrastive Regularization for Domain Generalization

    Daehee Kim, Seunghyun Park, Jinkyu Kim +1

    cs.CVcs.AIarXiv:2104.09841v12021
    Summaries:한국어
  37. DeSTSeg: Segmentation Guided Denoising Student-Teacher for Anomaly Detection

    Xuan Zhang, Shiyu Li, Xi Li +3

    cs.CVarXiv:2211.11317v22022
  38. BABEL: Bodies, Action and Behavior with English Labels

    Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou +2

    cs.CVcs.GRcs.LGarXiv:2106.09696v22021
  39. Fast Image Processing with Fully-Convolutional Networks

    Qifeng Chen, Jia Xu, Vladlen Koltun

    cs.CVcs.GRcs.LGarXiv:1709.00643v12017
  40. OminiControl: Minimal and Universal Control for Diffusion Transformer

    Zhenxiong Tan, Songhua Liu, Xingyi Yang +2

    cs.CVcs.AIcs.LGarXiv:2411.15098v62024
  41. Learning Joint Spatial-Temporal Transformations for Video Inpainting

    Yanhong Zeng, Jianlong Fu, Hongyang Chao

    cs.CVarXiv:2007.10247v12020
  42. FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular Cameras

    Anthony Hu, Zak Murez, Nikhil Mohan +5

    cs.CVcs.ROarXiv:2104.10490v32021
  43. DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization

    Jiahe Li, Jiawei Zhang, Xiao Bai +4

    cs.CVarXiv:2403.06912v32024
  44. Bi-Directional Cascade Network for Perceptual Edge Detection

    Jianzhong He, Shiliang Zhang, Ming Yang +2

    cs.CVarXiv:1902.10903v12019
  45. Distribution Matching Losses Can Hallucinate Features in Medical Image Translation

    Joseph Paul Cohen, Margaux Luck, Sina Honari

    cs.CVcs.LGarXiv:1805.08841v32018
  46. Deep Learning for Change Detection in Remote Sensing Images: Comprehensive Review and Meta-Analysis

    Lazhar Khelifi, Max Mignotte

    cs.CVarXiv:2006.05612v12020
  47. RGCNN: Regularized Graph CNN for Point Cloud Segmentation

    Gusi Te, Wei Hu, Zongming Guo +1

    cs.CVarXiv:1806.02952v12018
  48. Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks

    Anh Nguyen, Jason Yosinski, Jeff Clune

    cs.NEcs.CVarXiv:1602.03616v22016
  49. Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory

    Xizhou Zhu, Yuntao Chen, Hao Tian +10

    cs.AIcs.CLcs.CVarXiv:2305.17144v22023
  50. Adversarial Attacks and Defences Competition

    Alexey Kurakin, Ian Goodfellow, Samy Bengio +20

    cs.CVcs.CRcs.LGarXiv:1804.00097v12018
  51. A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large Datasets

    Bernhard Kerbl, Andréas Meuleman, Georgios Kopanas +3

    cs.CVcs.GRarXiv:2406.12080v12024
  52. UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated Content

    Zhengzhong Tu, Yilin Wang, Neil Birkbeck +2

    cs.CVeess.IVarXiv:2005.14354v22020
  53. Unleashing Text-to-Image Diffusion Models for Visual Perception

    Wenliang Zhao, Yongming Rao, Zuyan Liu +3

    cs.CVarXiv:2303.02153v12023
  54. Unsupervised Deep Homography: A Fast and Robust Homography Estimation Model

    Ty Nguyen, Steven W. Chen, Shreyas S. Shivakumar +2

    cs.CVarXiv:1709.03966v32017
  55. CoMatch: Semi-supervised Learning with Contrastive Graph Regularization

    Junnan Li, Caiming Xiong, Steven Hoi

    cs.LGcs.CVarXiv:2011.11183v22020
  56. gsplat: An Open-Source Library for Gaussian Splatting

    Vickie Ye, Ruilong Li, Justin Kerr +8

    cs.CVarXiv:2409.06765v12024
  57. Mining Cross-Image Semantics for Weakly Supervised Semantic Segmentation

    Guolei Sun, Wenguan Wang, Jifeng Dai +1

    cs.CVcs.LGeess.IVarXiv:2007.01947v22020
  58. Learning to Remember: A Synaptic Plasticity Driven Framework for Continual Learning

    Oleksiy Ostapenko, Mihai Puscas, Tassilo Klein +2

    cs.NEcs.CVcs.LGarXiv:1904.03137v42019
  59. A General Framework for Uncertainty Estimation in Deep Learning

    Antonio Loquercio, Mattia Segù, Davide Scaramuzza

    cs.CVstat.MLarXiv:1907.06890v42019
  60. Diversity Regularized Spatiotemporal Attention for Video-based Person Re-identification

    Shuang Li, Slawomir Bak, Peter Carr +1

    cs.CVarXiv:1803.09882v12018