Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,981 to 14,040 of 18,817
EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang +3
cs.CVcs.CLarXiv:2303.11331v22023Imbalance Problems in Object Detection: A Review
Kemal Oksuz, Baris Can Cam, Sinan Kalkan +1
cs.CVarXiv:1909.00169v32019Hyperspectral and Multispectral Image Fusion based on a Sparse Representation
Qi Wei, José Bioucas-Dias, Nicolas Dobigeon +1
cs.CVarXiv:1409.5729v12014OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
Qidong Huang, Xiaoyi Dong, Pan Zhang +6
cs.CVarXiv:2311.17911v32023Complexity Induction: Compositional Generalization via Structured Label Distortion
Aleksandr Abramov
cs.CVcs.AIcs.LGarXiv:2608.21464v12026Spatially Transformed Adversarial Examples
Chaowei Xiao, Jun-Yan Zhu, Bo Li +3
cs.CRcs.CVstat.MLarXiv:1801.02612v22018PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
Dan Xu, Wanli Ouyang, Xiaogang Wang +1
cs.CVarXiv:1805.04409v12018Simultaneously Localize, Segment and Rank the Camouflaged Objects
Yunqiu Lv, Jing Zhang, Yuchao Dai +4
cs.CVarXiv:2103.04011v22021ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding
Shuo Feng, Piji Li
cs.CVarXiv:2608.21878v12026Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos
Xiaoyang Liu, Kai Han
cs.CVcs.AIarXiv:2608.22102v12026Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models
Michael Jerge, Joseph Pelczar, Justin Downes
cs.CVarXiv:2608.21595v12026Learning to See by Moving
Pulkit Agrawal, Joao Carreira, Jitendra Malik
cs.CVcs.NEcs.ROarXiv:1505.01596v22015Measuring Gender Representation in Animated Films
David Bamman, Allison Cooper, Ruby Alvarez Rubio +2
cs.CVarXiv:2608.21429v12026BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing
Dongxu Li, Junnan Li, Steven C. H. Hoi
cs.CVcs.AIarXiv:2305.14720v22023AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation
Jinyi Zhou, Shuo Feng, Yufei Wu +1
cs.CVarXiv:2608.21926v12026Entity-Constrained CBCT Retrieval for Low-Resource Dental Record Completion
Nhi Ngoc-Yen Nguyen, Thai Nguyen, Kiet Huynh Cao Tuan +1
cs.CVarXiv:2608.21913v12026StereoDiffuer: Diffusion-based Progressive Geometry Modeling with Saliency Attention Perception for Stereo Matching
Bohan Li
cs.CVarXiv:2608.21710v12026Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou +3
cs.CVcs.AIarXiv:2608.19556v12026Channel-wise Autoregressive Entropy Models for Learned Image Compression
David Minnen, Saurabh Singh
eess.IVcs.CVcs.ITarXiv:2007.08739v12020DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning
Xiujin Liu, Tianyu Yang, Yilun Zhao +1
cs.CVcs.MAarXiv:2608.21438v12026UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning
Ruihao Li, Sen Wang, Zhiqiang Long +1
cs.CVarXiv:1709.06841v22017Unstructured Human Activity Detection from RGBD Images
Jaeyong Sung, Colin Ponce, Bart Selman +1
cs.ROcs.CVarXiv:1107.0169v22011Modeling Context Between Objects for Referring Expression Understanding
Varun K. Nagaraja, Vlad I. Morariu, Larry S. Davis
cs.CVarXiv:1608.00525v12016SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space
Jin Zhou, Hongliang Yang, Pengfei Xu +1
cs.CVarXiv:2608.21659v12026PAConv: Position Adaptive Convolution with Dynamic Kernel Assembling on Point Clouds
Mutian Xu, Runyu Ding, Hengshuang Zhao +1
cs.CVarXiv:2103.14635v22021A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series
Stanislas Chambon, Mathieu Galtier, Pierrick Arnal +2
stat.MLcs.CVq-bio.NCarXiv:1707.03321v22017Blended Latent Diffusion
Omri Avrahami, Ohad Fried, Dani Lischinski
cs.CVcs.GRcs.LGarXiv:2206.02779v22022FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows
Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz +5
cs.CVcs.AIarXiv:2608.21460v12026Large Margin Object Tracking with Circulant Feature Maps
Mengmeng Wang, Yong Liu, Zeyi Huang
cs.CVarXiv:1703.05020v22017BIMScript: Material-Aware Structured Scene Programs for BIM Ingestion
Prakash Kondibhau Naikade, Thomas B. Moeslund, Andreas Møgelmose
cs.CVarXiv:2608.21447v12026Transfer Learning with Deep Convolutional Neural Network (CNN) for Pneumonia Detection using Chest X-ray
Tawsifur Rahman, Muhammad E. H. Chowdhury, Amith Khandakar +5
eess.IVcs.CVcs.LGarXiv:2004.06578v12020Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge
Mingqi Gao, Sijie Li, Jungong Han
cs.CVarXiv:2608.22064v12026Towards Bitstream-corrupted Harsh Visual Understanding: Through Bitstream Language Modeling as Robust Semantic Priors
Chaoran Huang, Fangcheng Li, Tianyi Liu +2
cs.CVcs.MMarXiv:2608.21837v12026Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
Xin Dong, Shangyu Chen, Sinno Jialin Pan
cs.NEcs.CVcs.LGarXiv:1705.07565v22017When More References Hurt: Contamination-Aware DINOv2 Memory Banks for Few-Shot Steel Defect Detection
Hannaneh Kalantary, Javad Khoramdel
cs.CVcs.LGarXiv:2608.22082v22026Natural Language Object Retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach +3
cs.CVcs.CLarXiv:1511.04164v32015TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models
Chang Sun, Francesco Barbato, Matteo Caligiuri +1
cs.CVcs.AIarXiv:2608.21487v12026Automated Gleason Grading of Prostate Biopsies using Deep Learning
Wouter Bulten, Hans Pinckaers, Hester van Boven +6
eess.IVcs.CVarXiv:1907.07980v120193D Deep Learning on Medical Images: A Review
Satya P. Singh, Lipo Wang, Sukrit Gupta +3
q-bio.QMcs.CVcs.LGarXiv:2004.00218v42020Blind Super-Resolution With Iterative Kernel Correction
Jinjin Gu, Hannan Lu, Wangmeng Zuo +1
cs.CVarXiv:1904.03377v22019Hashing for Similarity Search: A Survey
Jingdong Wang, Heng Tao Shen, Jingkuan Song +1
cs.DScs.CVcs.DBarXiv:1408.2927v12014Eigen-CAM: Class Activation Map using Principal Components
Mohammed Bany Muhammad, Mohammed Yeasin
cs.CVcs.LGarXiv:2008.00299v12020Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models
Jinchang Zhu, Rong Fu, Yi Ding +3
cs.CVcs.CLarXiv:2608.21762v12026Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections
Elian Iluk, Gil Ben-Artzi
cs.LGcs.CVarXiv:2608.23794v12026Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition
Subhankar Chattoraj, Sawon Pratiher, Samiran Das +1
cs.CVarXiv:2608.21454v12026Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
Xue Yang, Xiaojiang Yang, Jirui Yang +4
cs.CVcs.AIcs.LGarXiv:2106.01883v520213D Human Pose Estimation = 2D Pose Estimation + Matching
Ching-Hang Chen, Deva Ramanan
cs.CVarXiv:1612.06524v22016NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
Tien-Ju Yang, Andrew Howard, Bo Chen +5
cs.CVarXiv:1804.03230v22018PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Jihyung Ko, Eunji Jung, Hyeongsub Kim +4
cs.CVcs.AIcs.CLarXiv:2608.21819v12026LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices
Riadul Islam, Joey Mule, Dhandeep Challagundla +3
cs.CVcs.AIeess.IVarXiv:2608.21764v12026Single-Image Depth Perception in the Wild
Weifeng Chen, Zhao Fu, Dawei Yang +1
cs.CVcs.AIarXiv:1604.03901v22016Searching Central Difference Convolutional Networks for Face Anti-Spoofing
Zitong Yu, Chenxu Zhao, Zezheng Wang +5
cs.CVarXiv:2003.04092v12020GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing
Sosmita Paul, Krishna Roy
cs.CVcs.LGarXiv:2608.22789v12026Material Recognition in the Wild with the Materials in Context Database
Sean Bell, Paul Upchurch, Noah Snavely +1
cs.CVarXiv:1412.0623v22014MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement
Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh +1
eess.IVcs.CVarXiv:2608.21495v12026What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
Jeonghun Baek, Geewook Kim, Junyeop Lee +5
cs.CVarXiv:1904.01906v42019A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
Seungho Eum, Unsang Park
cs.CVcs.HCarXiv:2608.23137v22026YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
Esteban Real, Jonathon Shlens, Stefano Mazzocchi +2
cs.CVarXiv:1702.00824v52017Data-free parameter pruning for Deep Neural Networks
Suraj Srinivas, R. Venkatesh Babu
cs.CVarXiv:1507.06149v12015Automatic Knee Osteoarthritis Diagnosis from Plain Radiographs: A Deep Learning-Based Approach
Aleksei Tiulpin, Jérôme Thevenot, Esa Rahtu +2
cs.CVarXiv:1710.10589v12017