Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,981 to 14,040 of 18,817

  1. EVA-02: A Visual Representation for Neon Genesis

    Yuxin Fang, Quan Sun, Xinggang Wang +3

    cs.CVcs.CLarXiv:2303.11331v22023
  2. Imbalance Problems in Object Detection: A Review

    Kemal Oksuz, Baris Can Cam, Sinan Kalkan +1

    cs.CVarXiv:1909.00169v32019
  3. Hyperspectral and Multispectral Image Fusion based on a Sparse Representation

    Qi Wei, José Bioucas-Dias, Nicolas Dobigeon +1

    cs.CVarXiv:1409.5729v12014
  4. OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

    Qidong Huang, Xiaoyi Dong, Pan Zhang +6

    cs.CVarXiv:2311.17911v32023
  5. Complexity Induction: Compositional Generalization via Structured Label Distortion

    Aleksandr Abramov

    cs.CVcs.AIcs.LGarXiv:2608.21464v12026
  6. Spatially Transformed Adversarial Examples

    Chaowei Xiao, Jun-Yan Zhu, Bo Li +3

    cs.CRcs.CVstat.MLarXiv:1801.02612v22018
  7. PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing

    Dan Xu, Wanli Ouyang, Xiaogang Wang +1

    cs.CVarXiv:1805.04409v12018
  8. Simultaneously Localize, Segment and Rank the Camouflaged Objects

    Yunqiu Lv, Jing Zhang, Yuchao Dai +4

    cs.CVarXiv:2103.04011v22021
  9. ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

    Shuo Feng, Piji Li

    cs.CVarXiv:2608.21878v12026
  10. Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

    Xiaoyang Liu, Kai Han

    cs.CVcs.AIarXiv:2608.22102v12026
  11. Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

    Michael Jerge, Joseph Pelczar, Justin Downes

    cs.CVarXiv:2608.21595v12026
  12. Learning to See by Moving

    Pulkit Agrawal, Joao Carreira, Jitendra Malik

    cs.CVcs.NEcs.ROarXiv:1505.01596v22015
  13. Measuring Gender Representation in Animated Films

    David Bamman, Allison Cooper, Ruby Alvarez Rubio +2

    cs.CVarXiv:2608.21429v12026
  14. BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing

    Dongxu Li, Junnan Li, Steven C. H. Hoi

    cs.CVcs.AIarXiv:2305.14720v22023
  15. AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

    Jinyi Zhou, Shuo Feng, Yufei Wu +1

    cs.CVarXiv:2608.21926v12026
  16. Entity-Constrained CBCT Retrieval for Low-Resource Dental Record Completion

    Nhi Ngoc-Yen Nguyen, Thai Nguyen, Kiet Huynh Cao Tuan +1

    cs.CVarXiv:2608.21913v12026
  17. StereoDiffuer: Diffusion-based Progressive Geometry Modeling with Saliency Attention Perception for Stereo Matching

    Bohan Li

    cs.CVarXiv:2608.21710v12026
  18. Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

    Yuanhao Ban, Jiaqi Feng, Hengguang Zhou +3

    cs.CVcs.AIarXiv:2608.19556v12026
  19. Channel-wise Autoregressive Entropy Models for Learned Image Compression

    David Minnen, Saurabh Singh

    eess.IVcs.CVcs.ITarXiv:2007.08739v12020
  20. DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

    Xiujin Liu, Tianyu Yang, Yilun Zhao +1

    cs.CVcs.MAarXiv:2608.21438v12026
  21. UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning

    Ruihao Li, Sen Wang, Zhiqiang Long +1

    cs.CVarXiv:1709.06841v22017
  22. Unstructured Human Activity Detection from RGBD Images

    Jaeyong Sung, Colin Ponce, Bart Selman +1

    cs.ROcs.CVarXiv:1107.0169v22011
  23. Modeling Context Between Objects for Referring Expression Understanding

    Varun K. Nagaraja, Vlad I. Morariu, Larry S. Davis

    cs.CVarXiv:1608.00525v12016
  24. SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space

    Jin Zhou, Hongliang Yang, Pengfei Xu +1

    cs.CVarXiv:2608.21659v12026
  25. PAConv: Position Adaptive Convolution with Dynamic Kernel Assembling on Point Clouds

    Mutian Xu, Runyu Ding, Hengshuang Zhao +1

    cs.CVarXiv:2103.14635v22021
  26. A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series

    Stanislas Chambon, Mathieu Galtier, Pierrick Arnal +2

    stat.MLcs.CVq-bio.NCarXiv:1707.03321v22017
  27. Blended Latent Diffusion

    Omri Avrahami, Ohad Fried, Dani Lischinski

    cs.CVcs.GRcs.LGarXiv:2206.02779v22022
  28. FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

    Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz +5

    cs.CVcs.AIarXiv:2608.21460v12026
  29. Large Margin Object Tracking with Circulant Feature Maps

    Mengmeng Wang, Yong Liu, Zeyi Huang

    cs.CVarXiv:1703.05020v22017
  30. BIMScript: Material-Aware Structured Scene Programs for BIM Ingestion

    Prakash Kondibhau Naikade, Thomas B. Moeslund, Andreas Møgelmose

    cs.CVarXiv:2608.21447v12026
  31. Transfer Learning with Deep Convolutional Neural Network (CNN) for Pneumonia Detection using Chest X-ray

    Tawsifur Rahman, Muhammad E. H. Chowdhury, Amith Khandakar +5

    eess.IVcs.CVcs.LGarXiv:2004.06578v12020
  32. Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge

    Mingqi Gao, Sijie Li, Jungong Han

    cs.CVarXiv:2608.22064v12026
  33. Towards Bitstream-corrupted Harsh Visual Understanding: Through Bitstream Language Modeling as Robust Semantic Priors

    Chaoran Huang, Fangcheng Li, Tianyi Liu +2

    cs.CVcs.MMarXiv:2608.21837v12026
  34. Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon

    Xin Dong, Shangyu Chen, Sinno Jialin Pan

    cs.NEcs.CVcs.LGarXiv:1705.07565v22017
  35. When More References Hurt: Contamination-Aware DINOv2 Memory Banks for Few-Shot Steel Defect Detection

    Hannaneh Kalantary, Javad Khoramdel

    cs.CVcs.LGarXiv:2608.22082v22026
  36. Natural Language Object Retrieval

    Ronghang Hu, Huazhe Xu, Marcus Rohrbach +3

    cs.CVcs.CLarXiv:1511.04164v32015
  37. TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

    Chang Sun, Francesco Barbato, Matteo Caligiuri +1

    cs.CVcs.AIarXiv:2608.21487v12026
  38. Automated Gleason Grading of Prostate Biopsies using Deep Learning

    Wouter Bulten, Hans Pinckaers, Hester van Boven +6

    eess.IVcs.CVarXiv:1907.07980v12019
  39. 3D Deep Learning on Medical Images: A Review

    Satya P. Singh, Lipo Wang, Sukrit Gupta +3

    q-bio.QMcs.CVcs.LGarXiv:2004.00218v42020
  40. Blind Super-Resolution With Iterative Kernel Correction

    Jinjin Gu, Hannan Lu, Wangmeng Zuo +1

    cs.CVarXiv:1904.03377v22019
  41. Hashing for Similarity Search: A Survey

    Jingdong Wang, Heng Tao Shen, Jingkuan Song +1

    cs.DScs.CVcs.DBarXiv:1408.2927v12014
  42. Eigen-CAM: Class Activation Map using Principal Components

    Mohammed Bany Muhammad, Mohammed Yeasin

    cs.CVcs.LGarXiv:2008.00299v12020
  43. Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models

    Jinchang Zhu, Rong Fu, Yi Ding +3

    cs.CVcs.CLarXiv:2608.21762v12026
  44. Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections

    Elian Iluk, Gil Ben-Artzi

    cs.LGcs.CVarXiv:2608.23794v12026
  45. Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition

    Subhankar Chattoraj, Sawon Pratiher, Samiran Das +1

    cs.CVarXiv:2608.21454v12026
  46. Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence

    Xue Yang, Xiaojiang Yang, Jirui Yang +4

    cs.CVcs.AIcs.LGarXiv:2106.01883v52021
  47. 3D Human Pose Estimation = 2D Pose Estimation + Matching

    Ching-Hang Chen, Deva Ramanan

    cs.CVarXiv:1612.06524v22016
  48. NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications

    Tien-Ju Yang, Andrew Howard, Bo Chen +5

    cs.CVarXiv:1804.03230v22018
  49. PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

    Jihyung Ko, Eunji Jung, Hyeongsub Kim +4

    cs.CVcs.AIcs.CLarXiv:2608.21819v12026
  50. LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

    Riadul Islam, Joey Mule, Dhandeep Challagundla +3

    cs.CVcs.AIeess.IVarXiv:2608.21764v12026
  51. Single-Image Depth Perception in the Wild

    Weifeng Chen, Zhao Fu, Dawei Yang +1

    cs.CVcs.AIarXiv:1604.03901v22016
  52. Searching Central Difference Convolutional Networks for Face Anti-Spoofing

    Zitong Yu, Chenxu Zhao, Zezheng Wang +5

    cs.CVarXiv:2003.04092v12020
  53. GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing

    Sosmita Paul, Krishna Roy

    cs.CVcs.LGarXiv:2608.22789v12026
  54. Material Recognition in the Wild with the Materials in Context Database

    Sean Bell, Paul Upchurch, Noah Snavely +1

    cs.CVarXiv:1412.0623v22014
  55. MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement

    Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh +1

    eess.IVcs.CVarXiv:2608.21495v12026
  56. What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis

    Jeonghun Baek, Geewook Kim, Junyeop Lee +5

    cs.CVarXiv:1904.01906v42019
  57. A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)

    Seungho Eum, Unsang Park

    cs.CVcs.HCarXiv:2608.23137v22026
  58. YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video

    Esteban Real, Jonathon Shlens, Stefano Mazzocchi +2

    cs.CVarXiv:1702.00824v52017
  59. Data-free parameter pruning for Deep Neural Networks

    Suraj Srinivas, R. Venkatesh Babu

    cs.CVarXiv:1507.06149v12015
  60. Automatic Knee Osteoarthritis Diagnosis from Plain Radiographs: A Deep Learning-Based Approach

    Aleksei Tiulpin, Jérôme Thevenot, Esa Rahtu +2

    cs.CVarXiv:1710.10589v12017