Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

8,881 to 8,940 of 18,867

  1. SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

    Amanuel Gizachew Abebe, Yasmin Moslem

    cs.CVcs.CLarXiv:2608.29974v12026
  2. Are Convolutional Neural Networks or Transformers more like human vision?

    Shikhar Tuli, Ishita Dasgupta, Erin Grant +1

    cs.CVarXiv:2105.07197v22021
  3. Invariant Attribute Profiles: A Spatial-Frequency Joint Feature Extractor for Hyperspectral Image Classification

    Danfeng Hong, Xin Wu, Pedram Ghamisi +3

    cs.CVarXiv:1912.08847v12019
  4. RS3Mamba: Visual State Space Model for Remote Sensing Images Semantic Segmentation

    Xianping Ma, Xiaokang Zhang, Man-On Pun

    cs.CVarXiv:2404.02457v12024
  5. Pushing the Envelope for RGB-based Dense 3D Hand Pose Estimation via Neural Rendering

    Seungryul Baek, Kwang In Kim, Tae-Kyun Kim

    cs.CVarXiv:1904.04196v22019
  6. SPViT: Enabling Faster Vision Transformers via Soft Token Pruning

    Zhenglun Kong, Peiyan Dong, Xiaolong Ma +9

    cs.CVcs.AIcs.ARarXiv:2112.13890v22021
  7. The Unsurprising Effectiveness of Pre-Trained Vision Models for Control

    Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam +1

    cs.CVcs.AIcs.LGarXiv:2203.03580v22022
  8. ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

    Xiaoqi Li, Mingxu Zhang, Yiran Geng +6

    cs.CVcs.ROarXiv:2312.16217v12023
  9. Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships

    Yong Liu, Ruiping Wang, Shiguang Shan +1

    cs.CVarXiv:1807.00119v12018
  10. NoScope: Optimizing Neural Network Queries over Video at Scale

    Daniel Kang, John Emmons, Firas Abuzaid +2

    cs.DBcs.CVarXiv:1703.02529v32017
  11. SlimYOLOv3: Narrower, Faster and Better for Real-Time UAV Applications

    Pengyi Zhang, Yunxin Zhong, Xiaoqiong Li

    cs.CVarXiv:1907.11093v12019
  12. EagerMOT: 3D Multi-Object Tracking via Sensor Fusion

    Aleksandr Kim, Aljoša Ošep, Laura Leal-Taixé

    cs.CVcs.ROarXiv:2104.14682v12021
  13. Foundation Models for Generalist Geospatial Artificial Intelligence

    Johannes Jakubik, Sujit Roy, C. E. Phillips +30

    cs.CVcs.LGarXiv:2310.18660v22023
  14. Deep Learning for Spacecraft Pose Estimation from Photorealistic Rendering

    Pedro F. Proenca, Yang Gao

    cs.CVcs.LGcs.ROarXiv:1907.04298v22019
  15. StyleDrop: Text-to-Image Generation in Any Style

    Kihyuk Sohn, Nataniel Ruiz, Kimin Lee +11

    cs.CVcs.AIarXiv:2306.00983v12023
  16. Few-Shot Adaptive Gaze Estimation

    Seonwook Park, Shalini De Mello, Pavlo Molchanov +3

    cs.CVarXiv:1905.01941v22019
  17. Numerical Coordinate Regression with Convolutional Neural Networks

    Aiden Nibali, Zhen He, Stuart Morgan +1

    cs.CVarXiv:1801.07372v22018
  18. SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos

    Zheyu Huang, Zijing Shi, Haozhe Luo +4

    cs.CLcs.CVarXiv:2608.30716v12026
  19. RAM: A Region-Aware Deep Model for Vehicle Re-Identification

    Xiaobin Liu, Shiliang Zhang, Qingming Huang +1

    cs.CVarXiv:1806.09283v12018
  20. EditGAN: High-Precision Semantic Image Editing

    Huan Ling, Karsten Kreis, Daiqing Li +3

    cs.CVcs.AIarXiv:2111.03186v12021
  21. Active Deep Learning for Classification of Hyperspectral Images

    Peng Liu, Hui Zhang, Kie B. Eom

    cs.LGcs.CVstat.MLarXiv:1611.10031v12016
  22. Self-training with progressive augmentation for unsupervised cross-domain person re-identification

    Xinyu Zhang, Jiewei Cao, Chunhua Shen +1

    cs.CVarXiv:1907.13315v12019
  23. SemEval-2020 Task 8: Memotion Analysis -- The Visuo-Lingual Metaphor!

    Chhavi Sharma, Deepesh Bhageria, William Scott +5

    cs.CVarXiv:2008.03781v12020
  24. Attention-based Ensemble for Deep Metric Learning

    Wonsik Kim, Bhavya Goyal, Kunal Chawla +2

    cs.CVarXiv:1804.00382v22018
  25. MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

    Shitao Tang, Fuyang Zhang, Jiacheng Chen +2

    cs.CVarXiv:2307.01097v72023
  26. Deep Equilibrium Architectures for Inverse Problems in Imaging

    Davis Gilton, Gregory Ongie, Rebecca Willett

    eess.IVcs.CVarXiv:2102.07944v22021
  27. Supervised Contrastive Replay: Revisiting the Nearest Class Mean Classifier in Online Class-Incremental Continual Learning

    Zheda Mai, Ruiwen Li, Hyunwoo Kim +1

    cs.LGcs.AIcs.CVarXiv:2103.13885v32021
  28. Self-Supervised Models are Continual Learners

    Enrico Fini, Victor G. Turrisi da Costa, Xavier Alameda-Pineda +3

    cs.CVcs.LGarXiv:2112.04215v22021
  29. Event-Based Fusion for Motion Deblurring with Cross-modal Attention

    Lei Sun, Christos Sakaridis, Jingyun Liang +6

    cs.CVarXiv:2112.00167v32021
  30. MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video

    Qunliang Xing, Zhenyu Guan, Mai Xu +3

    cs.CVcs.MMarXiv:1902.09707v62019
  31. Graph-based compression of dynamic 3D point cloud sequences

    Dorina Thanou, Philip A. Chou, Pascal Frossard

    cs.CVcs.GRarXiv:1506.06096v12015
  32. How to Fool Radiologists with Generative Adversarial Networks? A Visual Turing Test for Lung Cancer Diagnosis

    Maria J. M. Chuquicusma, Sarfaraz Hussein, Jeremy Burt +1

    cs.CVcs.AIcs.LGarXiv:1710.09762v22017
  33. Shape-IoU: More Accurate Metric considering Bounding Box Shape and Scale

    Hao Zhang, Shuaijie Zhang

    cs.CVarXiv:2312.17663v22023
  34. Improving Dermoscopic Image Segmentation with Enhanced Convolutional-Deconvolutional Networks

    Yading Yuan, Yeh-Chi Lo

    cs.CVarXiv:1709.09780v12017
  35. BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection

    Hedi Ben-younes, Rémi Cadene, Nicolas Thome +1

    cs.CVarXiv:1902.00038v22019
  36. Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving

    Xiaosong Jia, Penghao Wu, Li Chen +4

    cs.CVarXiv:2305.06242v12023
  37. Pushing the Boundaries of Boundary Detection using Deep Learning

    Iasonas Kokkinos

    cs.CVcs.LGarXiv:1511.07386v22015
  38. RIDI: Robust IMU Double Integration

    Hang Yan, Qi Shan, Yasutaka Furukawa

    cs.CVarXiv:1712.09004v22017
  39. ACTION-Net: Multipath Excitation for Action Recognition

    Zhengwei Wang, Qi She, Aljosa Smolic

    cs.CVarXiv:2103.07372v12021
  40. Scaling Robot Learning with Semantically Imagined Experience

    Tianhe Yu, Ted Xiao, Austin Stone +10

    cs.ROcs.AIcs.CLarXiv:2302.11550v12023
  41. DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model

    Yinghao Xu, Hao Tan, Fujun Luan +8

    cs.CVarXiv:2311.09217v12023
  42. Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models

    Kangwook Ko, Jaehyuk Jang, Wonjun Lee +2

    cs.CLcs.CVarXiv:2608.30649v12026
  43. Predicting Visual Features from Text for Image and Video Caption Retrieval

    Jianfeng Dong, Xirong Li, Cees G. M. Snoek

    cs.CVarXiv:1709.01362v32017
  44. GFF: Gated Fully Fusion for Semantic Segmentation

    Xiangtai Li, Houlong Zhao, Lei Han +2

    cs.CVarXiv:1904.01803v22019
  45. Generating Classification Weights with GNN Denoising Autoencoders for Few-Shot Learning

    Spyros Gidaris, Nikos Komodakis

    cs.CVcs.LGarXiv:1905.01102v12019
  46. Neural Image Compression for Gigapixel Histopathology Image Analysis

    David Tellez, Geert Litjens, Jeroen van der Laak +1

    cs.CVeess.IVarXiv:1811.02840v22018
  47. Stratified Transfer Learning for Cross-domain Activity Recognition

    Jindong Wang, Yiqiang Chen, Lisha Hu +2

    cs.CVcs.LGarXiv:1801.00820v12017
  48. Testing Deep Neural Networks

    Youcheng Sun, Xiaowei Huang, Daniel Kroening +3

    cs.LGcs.CVcs.SEarXiv:1803.04792v42018
  49. Auto-ReID: Searching for a Part-aware ConvNet for Person Re-Identification

    Ruijie Quan, Xuanyi Dong, Yu Wu +2

    cs.CVarXiv:1903.09776v42019
  50. Long-Term On-Board Prediction of People in Traffic Scenes under Uncertainty

    Apratim Bhattacharyya, Mario Fritz, Bernt Schiele

    cs.CVarXiv:1711.09026v22017
  51. Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

    Wanrong Zhu, Jack Hessel, Anas Awadalla +7

    cs.CVcs.CLarXiv:2304.06939v32023
  52. Domain Adaptation through Synthesis for Unsupervised Person Re-identification

    Slawomir Bak, Peter Carr, Jean-Francois Lalonde

    cs.CVarXiv:1804.10094v12018
  53. Actions ~ Transformations

    Xiaolong Wang, Ali Farhadi, Abhinav Gupta

    cs.CVarXiv:1512.00795v22015
  54. YOLO5Face: Why Reinventing a Face Detector

    Delong Qi, Weijun Tan, Qi Yao +1

    cs.CVarXiv:2105.12931v32021
  55. Semantic Video Segmentation by Gated Recurrent Flow Propagation

    David Nilsson, Cristian Sminchisescu

    cs.CVarXiv:1612.08871v22016
  56. Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild

    Dominik Kulon, Riza Alp Güler, Iasonas Kokkinos +2

    cs.CVarXiv:2004.01946v12020
  57. Superquadrics Revisited: Learning 3D Shape Parsing beyond Cuboids

    Despoina Paschalidou, Ali Osman Ulusoy, Andreas Geiger

    cs.CVarXiv:1904.09970v12019
  58. Preserving Semantic Relations for Zero-Shot Learning

    Yashas Annadani, Soma Biswas

    cs.CVarXiv:1803.03049v12018
  59. Gradient Descent Learns One-hidden-layer CNN: Don't be Afraid of Spurious Local Minima

    Simon S. Du, Jason D. Lee, Yuandong Tian +2

    cs.LGcs.AIcs.CVarXiv:1712.00779v22017
  60. Consistency-based Semi-supervised Active Learning: Towards Minimizing Labeling Cost

    Mingfei Gao, Zizhao Zhang, Guo Yu +3

    cs.LGcs.CVarXiv:1910.07153v22019