Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,341 to 14,400 of 18,815

  1. An Empirical Evaluation of Deep Learning on Highway Driving

    Brody Huval, Tao Wang, Sameep Tandon +10

    cs.ROcs.CVarXiv:1504.01716v32015
  2. DASNet: Dual attentive fully convolutional siamese networks for change detection of high resolution satellite images

    Jie Chen, Ziyang Yuan, Jian Peng +5

    cs.CVarXiv:2003.03608v22020
  3. COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

    Andreas Veit, Tomas Matera, Lukas Neumann +2

    cs.CVarXiv:1601.07140v22016
  4. Rethinking Rotated Object Detection with Gaussian Wasserstein Distance Loss

    Xue Yang, Junchi Yan, Qi Ming +3

    cs.CVcs.AIarXiv:2101.11952v42021
  5. Learning Video Object Segmentation from Static Images

    Anna Khoreva, Federico Perazzi, Rodrigo Benenson +2

    cs.CVarXiv:1612.02646v12016
  6. SemanticFusion: Dense 3D Semantic Mapping with Convolutional Neural Networks

    John McCormac, Ankur Handa, Andrew Davison +1

    cs.CVarXiv:1609.05130v22016
  7. LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS

    Zhiwen Fan, Kevin Wang, Kairun Wen +3

    cs.CVarXiv:2311.17245v62023
  8. Towards a Guideline for Evaluation Metrics in Medical Image Segmentation

    Dominik Müller, Iñaki Soto-Rey, Frank Kramer

    eess.IVcs.CVcs.LGarXiv:2202.05273v12022
  9. BundleFusion: Real-time Globally Consistent 3D Reconstruction using On-the-fly Surface Re-integration

    Angela Dai, Matthias Nießner, Michael Zollhöfer +2

    cs.GRcs.CVarXiv:1604.01093v32016
  10. Learning in the Frequency Domain

    Kai Xu, Minghai Qin, Fei Sun +3

    cs.CVarXiv:2002.12416v42020
  11. End-to-end Learning of Action Detection from Frame Glimpses in Videos

    Serena Yeung, Olga Russakovsky, Greg Mori +1

    cs.CVcs.LGarXiv:1511.06984v22015
  12. DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation

    Lukas Hoyer, Dengxin Dai, Luc Van Gool

    cs.CVarXiv:2111.14887v22021
  13. FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency Space

    Quande Liu, Cheng Chen, Jing Qin +2

    cs.CVarXiv:2103.06030v12021
  14. Generative and Discriminative Voxel Modeling with Convolutional Neural Networks

    Andrew Brock, Theodore Lim, J. M. Ritchie +1

    cs.CVcs.HCcs.LGarXiv:1608.04236v22016
  15. Cross Modal Distillation for Supervision Transfer

    Saurabh Gupta, Judy Hoffman, Jitendra Malik

    cs.CVarXiv:1507.00448v22015
  16. CODA-Prompt: COntinual Decomposed Attention-based Prompting for Rehearsal-Free Continual Learning

    James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta +6

    cs.CVcs.AIcs.LGarXiv:2211.13218v22022
  17. Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons

    Byeongho Heo, Minsik Lee, Sangdoo Yun +1

    cs.LGcs.CVstat.MLarXiv:1811.03233v22018
  18. Rethinking Classification and Localization for Object Detection

    Yue Wu, Yinpeng Chen, Lu Yuan +4

    cs.CVarXiv:1904.06493v42019
  19. PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

    Jiacong Xu, Zixiang Xiong, Shankar P. Bhattacharyya

    cs.CVcs.AIarXiv:2206.02066v32022
  20. Hyperbolic Hierarchical Clustering for Visual Representation Learning

    Jianan Wei, Guikun Chen, Zhiyuan Weng +3

    cs.CVcs.AIarXiv:2608.22665v12026
  21. Semantic Graph Convolutional Networks for 3D Human Pose Regression

    Long Zhao, Xi Peng, Yu Tian +2

    cs.CVarXiv:1904.03345v32019
  22. Adversarial Examples Improve Image Recognition

    Cihang Xie, Mingxing Tan, Boqing Gong +3

    cs.CVarXiv:1911.09665v22019
  23. Scaling Open-Vocabulary Image Segmentation with Image-Level Labels

    Golnaz Ghiasi, Xiuye Gu, Yin Cui +1

    cs.CVarXiv:2112.12143v22021
  24. Fine-Grained Head Pose Estimation Without Keypoints

    Nataniel Ruiz, Eunji Chong, James M. Rehg

    cs.CVarXiv:1710.00925v52017
  25. Towards Open World Object Detection

    K J Joseph, Salman Khan, Fahad Shahbaz Khan +1

    cs.CVcs.AIcs.LGarXiv:2103.02603v22021
  26. B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures

    Sebastián González, Karen Sanchez, José M. Saavedra +2

    cs.CVarXiv:2608.24364v12026
  27. Reliable Fidelity and Diversity Metrics for Generative Models

    Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh +2

    cs.CVcs.LGstat.MLarXiv:2002.09797v22020
  28. LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

    Ahmed Shehab Khan, Zhiyuan Li, Yan Tong

    cs.CVarXiv:2608.23880v12026
  29. Generative Face Completion

    Yijun Li, Sifei Liu, Jimei Yang +1

    cs.CVarXiv:1704.05838v12017
  30. DSOD: Learning Deeply Supervised Object Detectors from Scratch

    Zhiqiang Shen, Zhuang Liu, Jianguo Li +3

    cs.CVcs.LGarXiv:1708.01241v22017
  31. PolarMask: Single Shot Instance Segmentation with Polar Representation

    Enze Xie, Peize Sun, Xiaoge Song +4

    cs.CVarXiv:1909.13226v42019
  32. DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection

    Liming Jiang, Ren Li, Wayne Wu +2

    cs.CVcs.LGarXiv:2001.03024v22020
  33. ReGround-Surg: Reliability-Guided Anchor Grounding for Referring Surgical Video Segmentation

    Jiaxin Wen, Ming Yin, Lu Liu +1

    cs.CVarXiv:2608.24671v12026
  34. MaST: Motion-aware Sparse Pipeline for Lightweight Object Tracking

    Qingmao Wei, Fagui Liu, Dengke Zhang +2

    cs.CVarXiv:2608.24365v12026
  35. Beauty is in the ELBO of the Beholder: A Variational Account of Processing Fluency in Face Perception

    Francisco M. López, Jochen Triesch

    cs.CVq-bio.NCarXiv:2608.24219v12026
  36. COVID-CAPS: A Capsule Network-based Framework for Identification of COVID-19 cases from X-ray Images

    Parnian Afshar, Shahin Heidarian, Farnoosh Naderkhani +3

    cs.CVcs.LGeess.IVarXiv:2004.02696v22020
  37. Low-Rank Velocity Fields as a Structural Prior for Unsupervised 4D Medical Image Interpolation

    Haojin Li, Hengzhuo Wang, Chang Liu +3

    cs.CVarXiv:2608.24025v12026
  38. Deep Semantic Ranking Based Hashing for Multi-Label Image Retrieval

    Fang Zhao, Yongzhen Huang, Liang Wang +1

    cs.CVcs.LGarXiv:1501.06272v22015
  39. Velocity-coupled Representation Refinement for Satellite Orbit Prediction

    Yue Yang, Zhiqiang Wu, Saiyu Qi +1

    cs.CVarXiv:2608.23728v12026
  40. Open-Set Recognition: a Good Closed-Set Classifier is All You Need?

    Sagar Vaze, Kai Han, Andrea Vedaldi +1

    cs.CVcs.LGarXiv:2110.06207v22021
  41. MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

    Enxin Song, Wenhao Chai, Guanhong Wang +10

    cs.CVarXiv:2307.16449v42023
  42. SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

    Tianlv Huang, Hetian Guo, Ziyi Cai +6

    cs.CVcs.CLcs.GRarXiv:2608.24334v12026
  43. Land Use Classification in Remote Sensing Images by Convolutional Neural Networks

    Marco Castelluccio, Giovanni Poggi, Carlo Sansone +1

    cs.CVarXiv:1508.00092v12015
  44. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

    Feng Li, Renrui Zhang, Hao Zhang +5

    cs.CVcs.CLcs.LGarXiv:2407.07895v22024
  45. Moments in Time Dataset: one million videos for event understanding

    Mathew Monfort, Alex Andonian, Bolei Zhou +8

    cs.CVcs.AIarXiv:1801.03150v32018
  46. CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark

    Jiefeng Li, Can Wang, Hao Zhu +3

    cs.CVarXiv:1812.00324v22018
  47. A Strong Baseline and Batch Normalization Neck for Deep Person Re-identification

    Hao Luo, Wei Jiang, Youzhi Gu +4

    cs.CVarXiv:1906.08332v22019
  48. Why ReLU networks yield high-confidence predictions far away from the training data and how to mitigate the problem

    Matthias Hein, Maksym Andriushchenko, Julian Bitterwolf

    cs.LGcs.CVstat.MLarXiv:1812.05720v22018
  49. StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis

    Jiatao Gu, Lingjie Liu, Peng Wang +1

    cs.CVstat.MLarXiv:2110.08985v12021
  50. Markerless Pose Estimation for Resistance Training Technique Assessment

    Joseph Turner, Jeff Clark, Nawid Keshtmand

    cs.CVcs.AIarXiv:2608.24384v12026
  51. Survey of Nearest Neighbor Techniques

    Nitin Bhatia, Vandana

    cs.CVarXiv:1007.0085v12010
  52. Discrimination-aware Channel Pruning for Deep Neural Networks

    Zhuangwei Zhuang, Mingkui Tan, Bohan Zhuang +5

    cs.CVarXiv:1810.11809v32018
  53. Two-Stream Neural Networks for Tampered Face Detection

    Peng Zhou, Xintong Han, Vlad I. Morariu +1

    cs.CVarXiv:1803.11276v12018
  54. Image-Adaptive YOLO for Object Detection in Adverse Weather Conditions

    Wenyu Liu, Gaofeng Ren, Runsheng Yu +3

    cs.CVarXiv:2112.08088v32021
  55. Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training

    Hongkai Zhang, Hong Chang, Bingpeng Ma +2

    cs.CVarXiv:2004.06002v22020
  56. MUTAN: Multimodal Tucker Fusion for Visual Question Answering

    Hedi Ben-younes, Rémi Cadene, Matthieu Cord +1

    cs.CVarXiv:1705.06676v12017
  57. Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

    Boyuan Chen, Diego Marti Monso, Yilun Du +3

    cs.LGcs.CVcs.ROarXiv:2407.01392v42024
  58. Human Semantic Parsing for Person Re-identification

    Mahdi M. Kalayeh, Emrah Basaran, Muhittin Gokmen +2

    cs.CVarXiv:1804.00216v12018
  59. Generative Image Modeling using Style and Structure Adversarial Networks

    Xiaolong Wang, Abhinav Gupta

    cs.CVarXiv:1603.05631v22016
  60. TractSeg - Fast and accurate white matter tract segmentation

    Jakob Wasserthal, Peter Neher, Klaus H. Maier-Hein

    cs.CVeess.IVarXiv:1805.07103v22018