Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,241 to 3,300 of 18,839
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
Mingjie Pan, Jiaming Liu, Renrui Zhang +6
cs.CVarXiv:2309.09502v22023Transformer-based Image Compression
Ming Lu, Peiyao Guo, Huiqing Shi +2
eess.IVcs.CVarXiv:2111.06707v12021Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera
Weipeng Xu, Avishek Chatterjee, Michael Zollhoefer +4
cs.CVarXiv:1803.05959v22018Lattice Long Short-Term Memory for Human Action Recognition
Lin Sun, Kui Jia, Kevin Chen +3
cs.CVarXiv:1708.03958v12017MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
Jiaxu Wang, Yicheng Jiang, Tianlun He +8
cs.CVarXiv:2602.09878v22026EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
Shraman Pramanick, Yale Song, Sayan Nag +5
cs.CVarXiv:2307.05463v22023Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
cs.CVarXiv:2004.03686v32020Cross Modal Transformer: Towards Fast and Robust 3D Object Detection
Junjie Yan, Yingfei Liu, Jianjian Sun +4
cs.CVarXiv:2301.01283v32023VideoINR: Learning Video Implicit Neural Representation for Continuous Space-Time Super-Resolution
Zeyuan Chen, Yinbo Chen, Jingwen Liu +5
eess.IVcs.CVcs.LGarXiv:2206.04647v12022Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival Prediction
Guillaume Jaume, Anurag Vaidya, Richard Chen +3
cs.CVcs.AIq-bio.GNarXiv:2304.06819v22023RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs
Zhouxia Wang, Jiawei Zhang, Runjian Chen +2
cs.CVarXiv:2201.06374v32022AdaptiveWeighted Attention Network with Camera Spectral Sensitivity Prior for Spectral Reconstruction from RGB Images
Jiaojiao Li, Chaoxiong Wu, Rui Song +2
eess.IVcs.CVarXiv:2005.09305v12020Rethinking the Evaluation of Video Summaries
Mayu Otani, Yuta Nakashima, Esa Rahtu +1
cs.CVarXiv:1903.11328v22019HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features
Cheng Sun, Min Sun, Hwann-Tzong Chen
cs.CVarXiv:2011.11498v32020Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural Networks
Jesse Hagenaars, Federico Paredes-Vallés, Guido de Croon
cs.CVcs.AIcs.LGarXiv:2106.01862v22021Scale-Equivariant Steerable Networks
Ivan Sosnovik, Michał Szmaja, Arnold Smeulders
cs.CVcs.LGstat.MLarXiv:1910.11093v22019Appearance-Preserving 3D Convolution for Video-based Person Re-identification
Xinqian Gu, Hong Chang, Bingpeng Ma +2
cs.CVarXiv:2007.08434v22020Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation
Yicong Hong, Zun Wang, Qi Wu +1
cs.CVcs.CLcs.ROarXiv:2203.02764v12022SAT: 2D Semantics Assisted Training for 3D Visual Grounding
Zhengyuan Yang, Songyang Zhang, Liwei Wang +1
cs.CVarXiv:2105.11450v22021Zero-Shot Sketch-Image Hashing
Yuming Shen, Li Liu, Fumin Shen +1
cs.CVarXiv:1803.02284v12018Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective
Zeyuan Yin, Eric Xing, Zhiqiang Shen
cs.CVcs.AIcs.LGarXiv:2306.13092v32023MAT: Motion-Aware Multi-Object Tracking
Shoudong Han, Piao Huang, Hongwei Wang +4
cs.CVarXiv:2009.04794v22020GPRInvNet: Deep Learning-Based Ground Penetrating Radar Data Inversion for Tunnel Lining
Bin Liu, Yuxiao Ren, Hanchi Liu +4
cs.CVcs.LGeess.IVarXiv:1912.05759v32019Batch Normalization Embeddings for Deep Domain Generalization
Mattia Segu, Alessio Tonioni, Federico Tombari
cs.LGcs.CVarXiv:2011.12672v32020LongVLM: Efficient Long Video Understanding via Large Language Models
Yuetian Weng, Mingfei Han, Haoyu He +2
cs.CVarXiv:2404.03384v32024NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views
Dejia Xu, Yifan Jiang, Peihao Wang +3
cs.CVarXiv:2211.16431v22022A General Multi-Graph Matching Approach via Graduated Consistency-regularized Boosting
Junchi Yan, Minsu Cho, Hongyuan Zha +2
cs.CVarXiv:1502.05840v12015MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
Arjun Jain, Jonathan Tompson, Yann LeCun +1
cs.CVcs.LGcs.NEarXiv:1409.7963v12014An Architecture Combining Convolutional Neural Network (CNN) and Support Vector Machine (SVM) for Image Classification
Abien Fred Agarap
cs.CVcs.LGcs.NEarXiv:1712.03541v22017Domain-aware Visual Bias Eliminating for Generalized Zero-Shot Learning
Shaobo Min, Hantao Yao, Hongtao Xie +3
cs.CVarXiv:2003.13261v22020Scene Graph Generation: A Comprehensive Survey
Guangming Zhu, Liang Zhang, Youliang Jiang +8
cs.CVarXiv:2201.00443v22022CC-4DGS: Computational Deformation and Point-Cloud Compression for Storage-Efficient Dynamic Gaussian Splatting
Kyungdae Park, Chae Eun Rhee
cs.CVarXiv:2609.02184v12026Training-Free Speech-Centric Omni Understanding with Frozen VLMs
Ankan Deria, Hanoona Rasheed, Xilin He +2
eess.AScs.CVcs.SDarXiv:2609.04242v12026Collaborative On-Sensor Array Cameras
Jipeng Sun, Kaixuan Wei, Thomas Eboli +6
physics.opticscs.CVcs.DCarXiv:2506.04061v12025Heterogeneous Knowledge Distillation using Information Flow Modeling
Nikolaos Passalis, Maria Tzelepi, Anastasios Tefas
cs.CVarXiv:2005.00727v12020GridMM: Grid Memory Map for Vision-and-Language Navigation
Zihan Wang, Xiangyang Li, Jiahao Yang +2
cs.CVcs.AIarXiv:2307.12907v42023Meta-Tracker: Fast and Robust Online Adaptation for Visual Object Trackers
Eunbyung Park, Alexander C. Berg
cs.CVcs.LGarXiv:1801.03049v22018VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck
Feiran Zhang, Yixin Wu, Zhenghua Wang +4
cs.CVcs.AIarXiv:2601.05547v22026A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models
Woojeong Jin, Yu Cheng, Yelong Shen +2
cs.CVcs.CLarXiv:2110.08484v22021Bridging Category-level and Instance-level Semantic Image Segmentation
Zifeng Wu, Chunhua Shen, Anton van den Hengel
cs.CVarXiv:1605.06885v12016Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation
Xian Liu, Yinghao Xu, Qianyi Wu +3
cs.CVcs.GRcs.LGarXiv:2201.07786v12022SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds
Pei Sun, Mingxing Tan, Weiyue Wang +4
cs.CVarXiv:2210.07372v12022VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Jiannan Wu, Muyan Zhong, Sen Xing +10
cs.CVarXiv:2406.08394v32024Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
Vivek Chavan, Yahuan Shi, Oliver Heimann +2
cs.ROcs.CVarXiv:2609.05369v12026Image Difference Quantification Using Autoencoder-Based Latent Representations
Manish Sharma, Timothy Yim, Clifton Forlines
cs.CVarXiv:2608.24782v12026GRASS: Generative Recursive Autoencoders for Shape Structures
Jun Li, Kai Xu, Siddhartha Chaudhuri +3
cs.GRcs.CVarXiv:1705.02090v22017Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
Sen Wang, Bangwei Liu, Zhenkun Gao +4
cs.AIcs.CVarXiv:2601.10744v22026Real-World Multi-Modal and Longitudinal Lung Cancer Dataset
Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago +1
eess.IVcs.CVarXiv:2609.05202v12026Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery
Nazim-E-Alam
eess.IVcs.CVarXiv:2609.05140v12026Long-Short Transformer: Efficient Transformers for Language and Vision
Chen Zhu, Wei Ping, Chaowei Xiao +4
cs.CVcs.CLcs.LGarXiv:2107.02192v32021Using DUCK-Net for Polyp Image Segmentation
Razvan-Gabriel Dumitru, Darius Peteleaza, Catalin Craciun
cs.CVcs.LGarXiv:2311.02239v12023CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving
Zhe Huang, Zhaoxin Fan, Shuo Wang +3
cs.ROcs.CVarXiv:2609.04807v12026BEAM3R: Beam's-eye-view architecture with Mamba-3 for implicit dose reconstruction
Chen Cheng, Michael Ferraro, James Grover +2
physics.med-phcs.CVarXiv:2609.04747v12026Joint Generative and Contrastive Learning for Unsupervised Person Re-identification
Hao Chen, Yaohui Wang, Benoit Lagadec +2
cs.CVarXiv:2012.09071v22020Learning Spatial-Spectral Refinement and Calibrating Complementary Observations for Hyperspectral Image Super-Resolution
Liqian Yang, Xingchi Chen, Xinfeng Gui +2
cs.CVarXiv:2609.05303v12026RMDL: Random Multimodel Deep Learning for Classification
Kamran Kowsari, Mojtaba Heidarysafa, Donald E. Brown +2
cs.LGcs.AIcs.CVarXiv:1805.01890v22018Quantum Adversarial Machine Learning
Sirui Lu, Lu-Ming Duan, Dong-Ling Deng
quant-phcond-mat.dis-nncond-mat.str-elarXiv:2001.00030v12019Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents
Daniel Davila, Ravikumar Balakrishnan, Mike Cochran
cs.CVarXiv:2609.05239v12026Few-Shot Learning with Localization in Realistic Settings
Davis Wertheimer, Bharath Hariharan
cs.CVcs.AIcs.LGarXiv:1904.08502v22019From Interpretability Methods to Interpretable Models
Julien Colin, Nuria Oliver, Thomas Serre
cs.CVcs.HCarXiv:2609.05399v12026