Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,241 to 3,300 of 18,839

  1. RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision

    Mingjie Pan, Jiaming Liu, Renrui Zhang +6

    cs.CVarXiv:2309.09502v22023
  2. Transformer-based Image Compression

    Ming Lu, Peiyao Guo, Huiqing Shi +2

    eess.IVcs.CVarXiv:2111.06707v12021
  3. Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera

    Weipeng Xu, Avishek Chatterjee, Michael Zollhoefer +4

    cs.CVarXiv:1803.05959v22018
  4. Lattice Long Short-Term Memory for Human Action Recognition

    Lin Sun, Kui Jia, Kevin Chen +3

    cs.CVarXiv:1708.03958v12017
  5. MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

    Jiaxu Wang, Yicheng Jiang, Tianlun He +8

    cs.CVarXiv:2602.09878v22026
  6. EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

    Shraman Pramanick, Yale Song, Sayan Nag +5

    cs.CVarXiv:2307.05463v22023
  7. Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation

    Hanbyul Joo, Natalia Neverova, Andrea Vedaldi

    cs.CVarXiv:2004.03686v32020
  8. Cross Modal Transformer: Towards Fast and Robust 3D Object Detection

    Junjie Yan, Yingfei Liu, Jianjian Sun +4

    cs.CVarXiv:2301.01283v32023
  9. VideoINR: Learning Video Implicit Neural Representation for Continuous Space-Time Super-Resolution

    Zeyuan Chen, Yinbo Chen, Jingwen Liu +5

    eess.IVcs.CVcs.LGarXiv:2206.04647v12022
  10. Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival Prediction

    Guillaume Jaume, Anurag Vaidya, Richard Chen +3

    cs.CVcs.AIq-bio.GNarXiv:2304.06819v22023
  11. RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs

    Zhouxia Wang, Jiawei Zhang, Runjian Chen +2

    cs.CVarXiv:2201.06374v32022
  12. AdaptiveWeighted Attention Network with Camera Spectral Sensitivity Prior for Spectral Reconstruction from RGB Images

    Jiaojiao Li, Chaoxiong Wu, Rui Song +2

    eess.IVcs.CVarXiv:2005.09305v12020
  13. Rethinking the Evaluation of Video Summaries

    Mayu Otani, Yuta Nakashima, Esa Rahtu +1

    cs.CVarXiv:1903.11328v22019
  14. HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features

    Cheng Sun, Min Sun, Hwann-Tzong Chen

    cs.CVarXiv:2011.11498v32020
  15. Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural Networks

    Jesse Hagenaars, Federico Paredes-Vallés, Guido de Croon

    cs.CVcs.AIcs.LGarXiv:2106.01862v22021
  16. Scale-Equivariant Steerable Networks

    Ivan Sosnovik, Michał Szmaja, Arnold Smeulders

    cs.CVcs.LGstat.MLarXiv:1910.11093v22019
  17. Appearance-Preserving 3D Convolution for Video-based Person Re-identification

    Xinqian Gu, Hong Chang, Bingpeng Ma +2

    cs.CVarXiv:2007.08434v22020
  18. Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

    Yicong Hong, Zun Wang, Qi Wu +1

    cs.CVcs.CLcs.ROarXiv:2203.02764v12022
  19. SAT: 2D Semantics Assisted Training for 3D Visual Grounding

    Zhengyuan Yang, Songyang Zhang, Liwei Wang +1

    cs.CVarXiv:2105.11450v22021
  20. Zero-Shot Sketch-Image Hashing

    Yuming Shen, Li Liu, Fumin Shen +1

    cs.CVarXiv:1803.02284v12018
  21. Squeeze, Recover and Relabel: Dataset Condensation at ImageNet Scale From A New Perspective

    Zeyuan Yin, Eric Xing, Zhiqiang Shen

    cs.CVcs.AIcs.LGarXiv:2306.13092v32023
  22. MAT: Motion-Aware Multi-Object Tracking

    Shoudong Han, Piao Huang, Hongwei Wang +4

    cs.CVarXiv:2009.04794v22020
  23. GPRInvNet: Deep Learning-Based Ground Penetrating Radar Data Inversion for Tunnel Lining

    Bin Liu, Yuxiao Ren, Hanchi Liu +4

    cs.CVcs.LGeess.IVarXiv:1912.05759v32019
  24. Batch Normalization Embeddings for Deep Domain Generalization

    Mattia Segu, Alessio Tonioni, Federico Tombari

    cs.LGcs.CVarXiv:2011.12672v32020
  25. LongVLM: Efficient Long Video Understanding via Large Language Models

    Yuetian Weng, Mingfei Han, Haoyu He +2

    cs.CVarXiv:2404.03384v32024
  26. NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views

    Dejia Xu, Yifan Jiang, Peihao Wang +3

    cs.CVarXiv:2211.16431v22022
  27. A General Multi-Graph Matching Approach via Graduated Consistency-regularized Boosting

    Junchi Yan, Minsu Cho, Hongyuan Zha +2

    cs.CVarXiv:1502.05840v12015
  28. MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation

    Arjun Jain, Jonathan Tompson, Yann LeCun +1

    cs.CVcs.LGcs.NEarXiv:1409.7963v12014
  29. An Architecture Combining Convolutional Neural Network (CNN) and Support Vector Machine (SVM) for Image Classification

    Abien Fred Agarap

    cs.CVcs.LGcs.NEarXiv:1712.03541v22017
  30. Domain-aware Visual Bias Eliminating for Generalized Zero-Shot Learning

    Shaobo Min, Hantao Yao, Hongtao Xie +3

    cs.CVarXiv:2003.13261v22020
  31. Scene Graph Generation: A Comprehensive Survey

    Guangming Zhu, Liang Zhang, Youliang Jiang +8

    cs.CVarXiv:2201.00443v22022
  32. CC-4DGS: Computational Deformation and Point-Cloud Compression for Storage-Efficient Dynamic Gaussian Splatting

    Kyungdae Park, Chae Eun Rhee

    cs.CVarXiv:2609.02184v12026
  33. Training-Free Speech-Centric Omni Understanding with Frozen VLMs

    Ankan Deria, Hanoona Rasheed, Xilin He +2

    eess.AScs.CVcs.SDarXiv:2609.04242v12026
  34. Collaborative On-Sensor Array Cameras

    Jipeng Sun, Kaixuan Wei, Thomas Eboli +6

    physics.opticscs.CVcs.DCarXiv:2506.04061v12025
  35. Heterogeneous Knowledge Distillation using Information Flow Modeling

    Nikolaos Passalis, Maria Tzelepi, Anastasios Tefas

    cs.CVarXiv:2005.00727v12020
  36. GridMM: Grid Memory Map for Vision-and-Language Navigation

    Zihan Wang, Xiangyang Li, Jiahao Yang +2

    cs.CVcs.AIarXiv:2307.12907v42023
  37. Meta-Tracker: Fast and Robust Online Adaptation for Visual Object Trackers

    Eunbyung Park, Alexander C. Berg

    cs.CVcs.LGarXiv:1801.03049v22018
  38. VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck

    Feiran Zhang, Yixin Wu, Zhenghua Wang +4

    cs.CVcs.AIarXiv:2601.05547v22026
  39. A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

    Woojeong Jin, Yu Cheng, Yelong Shen +2

    cs.CVcs.CLarXiv:2110.08484v22021
  40. Bridging Category-level and Instance-level Semantic Image Segmentation

    Zifeng Wu, Chunhua Shen, Anton van den Hengel

    cs.CVarXiv:1605.06885v12016
  41. Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation

    Xian Liu, Yinghao Xu, Qianyi Wu +3

    cs.CVcs.GRcs.LGarXiv:2201.07786v12022
  42. SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds

    Pei Sun, Mingxing Tan, Weiyue Wang +4

    cs.CVarXiv:2210.07372v12022
  43. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

    Jiannan Wu, Muyan Zhong, Sen Xing +10

    cs.CVarXiv:2406.08394v32024
  44. Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

    Vivek Chavan, Yahuan Shi, Oliver Heimann +2

    cs.ROcs.CVarXiv:2609.05369v12026
  45. Image Difference Quantification Using Autoencoder-Based Latent Representations

    Manish Sharma, Timothy Yim, Clifton Forlines

    cs.CVarXiv:2608.24782v12026
  46. GRASS: Generative Recursive Autoencoders for Shape Structures

    Jun Li, Kai Xu, Siddhartha Chaudhuri +3

    cs.GRcs.CVarXiv:1705.02090v22017
  47. Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

    Sen Wang, Bangwei Liu, Zhenkun Gao +4

    cs.AIcs.CVarXiv:2601.10744v22026
  48. Real-World Multi-Modal and Longitudinal Lung Cancer Dataset

    Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago +1

    eess.IVcs.CVarXiv:2609.05202v12026
  49. Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery

    Nazim-E-Alam

    eess.IVcs.CVarXiv:2609.05140v12026
  50. Long-Short Transformer: Efficient Transformers for Language and Vision

    Chen Zhu, Wei Ping, Chaowei Xiao +4

    cs.CVcs.CLcs.LGarXiv:2107.02192v32021
  51. Using DUCK-Net for Polyp Image Segmentation

    Razvan-Gabriel Dumitru, Darius Peteleaza, Catalin Craciun

    cs.CVcs.LGarXiv:2311.02239v12023
  52. CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving

    Zhe Huang, Zhaoxin Fan, Shuo Wang +3

    cs.ROcs.CVarXiv:2609.04807v12026
  53. BEAM3R: Beam's-eye-view architecture with Mamba-3 for implicit dose reconstruction

    Chen Cheng, Michael Ferraro, James Grover +2

    physics.med-phcs.CVarXiv:2609.04747v12026
  54. Joint Generative and Contrastive Learning for Unsupervised Person Re-identification

    Hao Chen, Yaohui Wang, Benoit Lagadec +2

    cs.CVarXiv:2012.09071v22020
  55. Learning Spatial-Spectral Refinement and Calibrating Complementary Observations for Hyperspectral Image Super-Resolution

    Liqian Yang, Xingchi Chen, Xinfeng Gui +2

    cs.CVarXiv:2609.05303v12026
  56. RMDL: Random Multimodel Deep Learning for Classification

    Kamran Kowsari, Mojtaba Heidarysafa, Donald E. Brown +2

    cs.LGcs.AIcs.CVarXiv:1805.01890v22018
  57. Quantum Adversarial Machine Learning

    Sirui Lu, Lu-Ming Duan, Dong-Ling Deng

    quant-phcond-mat.dis-nncond-mat.str-elarXiv:2001.00030v12019
  58. Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents

    Daniel Davila, Ravikumar Balakrishnan, Mike Cochran

    cs.CVarXiv:2609.05239v12026
  59. Few-Shot Learning with Localization in Realistic Settings

    Davis Wertheimer, Bharath Hariharan

    cs.CVcs.AIcs.LGarXiv:1904.08502v22019
  60. From Interpretability Methods to Interpretable Models

    Julien Colin, Nuria Oliver, Thomas Serre

    cs.CVcs.HCarXiv:2609.05399v12026