Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,341 to 2,400 of 18,786
SegSort: Segmentation by Discriminative Sorting of Segments
Jyh-Jing Hwang, Stella X. Yu, Jianbo Shi +4
cs.CVcs.LGeess.IVarXiv:1910.06962v22019XIRL: Cross-embodiment Inverse Reinforcement Learning
Kevin Zakka, Andy Zeng, Pete Florence +3
cs.ROcs.AIcs.CVarXiv:2106.03911v32021SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection
Yichen Xie, Chenfeng Xu, Marie-Julie Rakotosaona +5
cs.CVarXiv:2304.14340v12023Comparative Study of Anatomical and Learned Features in AI Models for Structural Brain MRI
Boyang Yu, Miquel Lopez Escoriza, Long Chen +3
cs.CVcs.AIarXiv:2609.06807v12026CompGS: Smaller and Faster Gaussian Splatting with Vector Quantization
KL Navaneet, Kossar Pourahmadi Meibodi, Soroush Abbasi Koohpayegani +1
cs.CVarXiv:2311.18159v32023Blended Multi-Modal Deep ConvNet Features for Diabetic Retinopathy Severity Prediction
J. D. Bodapati, N. Veeranjaneyulu, S. N. Shareef +4
eess.IVcs.CVcs.LGarXiv:2006.00197v12020Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
Bohan Zhuang, Chunhua Shen, Mingkui Tan +2
cs.CVarXiv:1811.10413v22018Mumford-Shah Loss Functional for Image Segmentation with Deep Learning
Boah Kim, Jong Chul Ye
cs.CVcs.LGstat.MLarXiv:1904.02872v22019Learning to Reconstruct Shapes from Unseen Classes
Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang +3
cs.CVcs.AIarXiv:1812.11166v12018When 3D Gaussian Splatting Recovers Real Surfaces
Songhe Wang, David Johnathan Miller
cs.LGcs.CVarXiv:2608.30054v12026Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation
Runzhong Zhang, Yueqi Duan, Yang Chen +4
cs.CVarXiv:2609.08167v12026se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains
Bowen Wen, Chaitanya Mitash, Baozhang Ren +1
cs.CVcs.GRcs.LGarXiv:2007.13866v12020Unified Quality Assessment of In-the-Wild Videos with Mixed Datasets Training
Dingquan Li, Tingting Jiang, Ming Jiang
cs.CVcs.MMeess.IVarXiv:2011.04263v22020Selective Refinement Network for High Performance Face Detection
Cheng Chi, Shifeng Zhang, Junliang Xing +3
cs.CVarXiv:1809.02693v12018Identity-Preserving Text-to-Video Generation by Frequency Decomposition
Shenghai Yuan, Jinfa Huang, Xianyi He +5
cs.CVcs.MMarXiv:2411.17440v32024Pyramid Diffusion Models For Low-light Image Enhancement
Dewei Zhou, Zongxin Yang, Yi Yang
cs.CVarXiv:2305.10028v12023FSAN: Flow State Attention Network for Aerodynamic Prediction
Wenxuan Jin, Jianguo Yao, Haibing Guan +1
cs.CVcs.AIarXiv:2609.06660v12026ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics
Chin Ting Hsu, Yu-Syuan Xu, Ling Zou +2
cs.CLcs.AIcs.CVarXiv:2609.06663v12026Deep Regression Forests for Age Estimation
Wei Shen, Yilu Guo, Yan Wang +3
cs.CVarXiv:1712.07195v12017Debiased Learning from Naturally Imbalanced Pseudo-Labels
Xudong Wang, Zhirong Wu, Long Lian +1
cs.LGcs.CLcs.CVarXiv:2201.01490v22022Robust Training under Label Noise by Over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu +1
cs.LGcs.AIcs.CVarXiv:2202.14026v22022Unsupervised Meta-Learning For Few-Shot Image Classification
Siavash Khodadadeh, Ladislau Bölöni, Mubarak Shah
cs.CVcs.LGarXiv:1811.11819v22018Deep Sinogram Completion with Image Prior for Metal Artifact Reduction in CT Images
Lequan Yu, Zhicheng Zhang, Xiaomeng Li +1
eess.IVcs.CVarXiv:2009.07469v12020Dictionary Learning and Sparse Coding on Grassmann Manifolds: An Extrinsic Solution
Mehrtash Harandi, Conrad Sanderson, Chunhua Shen +1
cs.CVarXiv:1310.4891v12013Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images
Aron Yu, Kristen Grauman
cs.CVarXiv:1612.06341v22016Companion-style QA Assistance in Ego-Vision
Hangyu Qin, Junbin Xiao, Shenglang Zhang +1
cs.CVcs.AIarXiv:2609.06721v12026Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models
Bayar Menzat, Maximilian Süss, Ruizhi Wang +3
cs.CVcs.AIcs.CLarXiv:2609.06704v12026FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval
Dehong Gao, Linbo Jin, Ben Chen +5
cs.IRcs.CVcs.LGarXiv:2005.09801v22020Attention-Enhanced Deep Features with Heterogeneous Ensemble Learning for Glaucoma Detection
Abdullah Al Shafi, Nishat Sadaf Lira, Abrar Hasan +2
cs.CVcs.AIcs.LGarXiv:2609.06699v12026ConvMAE: Masked Convolution Meets Masked Autoencoders
Peng Gao, Teli Ma, Hongsheng Li +3
cs.CVarXiv:2205.03892v22022Towards Geospatial Foundation Models via Continual Pretraining
Matias Mendieta, Boran Han, Xingjian Shi +2
cs.CVarXiv:2302.04476v32023Overcoming Catastrophic Forgetting in Incremental Object Detection via Elastic Response Distillation
Tao Feng, Mang Wang, Hangjie Yuan
cs.CVarXiv:2204.02136v12022Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN
Zhenwei He, Lei Zhang
cs.CVarXiv:2007.01571v12020Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training
Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel
cs.CVcs.AIcs.HCarXiv:1812.06145v22018Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse Coding
David Klindt, Lukas Schott, Yash Sharma +4
stat.MLcs.CVcs.LGarXiv:2007.10930v22020Deep Video Generation, Prediction and Completion of Human Action Sequences
Haoye Cai, Chunyan Bai, Yu-Wing Tai +1
cs.CVstat.MLarXiv:1711.08682v32017A Unified Continual Learning Framework with General Parameter-Efficient Tuning
Qiankun Gao, Chen Zhao, Yifan Sun +4
cs.CVarXiv:2303.10070v22023Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation
Jichang Li, Guanbin Li, Yemin Shi +1
cs.CVarXiv:2104.09415v120213D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
Yanqi Bao, Tianyu Ding, Jing Huo +5
cs.CVarXiv:2407.17418v22024Gray Level Co-Occurrence Matrices: Generalisation and Some New Features
Bino Sebastian, A. Unnikrishnan, Kannan Balakrishnan
cs.CVarXiv:1205.4831v12012Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
Jiayu Wang, Yifei Ming, Zhenmei Shi +4
cs.CVcs.AIarXiv:2406.14852v22024DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
Wanli Ouyang, Ping Luo, Xingyu Zeng +12
cs.CVarXiv:1409.3505v12014Few-Shot Defect Image Generation via Defect-Aware Feature Manipulation
Yuxuan Duan, Yan Hong, Li Niu +1
cs.CVarXiv:2303.02389v12023Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds
Jiacheng Wei, Guosheng Lin, Kim-Hui Yap +2
cs.CVarXiv:2003.13035v12020TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang +6
cs.CVarXiv:2012.04638v12020Accel: A Corrective Fusion Network for Efficient Semantic Segmentation on Video
Samvit Jain, Xin Wang, Joseph Gonzalez
cs.CVcs.LGarXiv:1807.06667v42018When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization
Eyal Hanania, Daniel Arkushin, Naveh Ayal +4
cs.CVcs.AIarXiv:2609.06646v12026Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification
Eduardo H. P. Pooch, Pedro L. Ballester, Rodrigo C. Barros
eess.IVcs.AIcs.CVarXiv:1909.01940v22019Diffusion Models, Image Super-Resolution And Everything: A Survey
Brian B. Moser, Arundhati S. Shanbhag, Federico Raue +3
cs.CVcs.AIcs.LGarXiv:2401.00736v32024Audio Surveillance: a Systematic Review
Marco Crocco, Marco Cristani, Andrea Trucco +1
cs.SDcs.CVcs.MMarXiv:1409.7787v12014Evaluating the Single-Shot MultiBox Detector and YOLO Deep Learning Models for the Detection of Tomatoes in a Greenhouse
Sandro A. Magalhães, Luís Castro, Germano Moreira +4
cs.CVcs.ROarXiv:2109.00810v12021Few-Example Object Detection with Model Communication
Xuanyi Dong, Liang Zheng, Fan Ma +2
cs.CVarXiv:1706.08249v82017Intra-Retinal Layer Segmentation of 3D Optical Coherence Tomography Using Coarse Grained Diffusion Map
Raheleh Kafieh, Hossein Rabbani, Michael D. Abramoff +1
cs.CVarXiv:1210.0310v22012Visual Explanations From Deep 3D Convolutional Neural Networks for Alzheimer's Disease Classification
Chengliang Yang, Anand Rangarajan, Sanjay Ranka
cs.CVcs.AIcs.LGarXiv:1803.02544v32018CoreDiff: Contextual Error-Modulated Generalized Diffusion Model for Low-Dose CT Denoising and Generalization
Qi Gao, Zilong Li, Junping Zhang +2
eess.IVcs.CVcs.LGarXiv:2304.01814v22023On the generalization of GAN image forensics
Xinsheng Xuan, Bo Peng, Wei Wang +1
cs.CVcs.LGstat.MLarXiv:1902.11153v22019Machine Vision for Natural Gas Methane Emissions Detection Using an Infrared Camera
Jingfan Wang, Lyne P. Tchapmi, Arvind P. Ravikumara +5
cs.CVcs.LGeess.IVarXiv:1904.08500v12019In-context learning enables multimodal large language models to classify cancer pathology images
Dyke Ferber, Georg Wölflein, Isabella C. Wiest +8
cs.CVarXiv:2403.07407v12024Deep Learning-Based Autonomous Driving Systems: A Survey of Attacks and Defenses
Yao Deng, Tiehua Zhang, Guannan Lou +3
cs.LGcs.CRcs.CVarXiv:2104.01789v22021Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier
Jingtao Lei, Hongji Li, Dexiang Shu
cs.LGcs.AIcs.CVarXiv:2609.06590v12026