Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
601 to 660 of 18,855
Robust Learning Through Cross-Task Consistency
Amir Zamir, Alexander Sax, Teresa Yeo +6
cs.CVcs.GRcs.LGarXiv:2006.04096v12020AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Qifan Yu, Wei Chow, Zhongqi Yue +7
cs.CVarXiv:2411.15738v32024CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends
Seojin Ji, Yoojin Kwon, Hyung-Sin Kim
cs.CVarXiv:2609.04784v12026A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
Vladimir Iashin, Esa Rahtu
cs.CVcs.CLcs.LGarXiv:2005.08271v22020Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
Sixu Yan, Shikang Wang, Binhua Huang +11
cs.ROcs.AIcs.CVarXiv:2609.04096v12026SoundNet: Learning Sound Representations from Unlabeled Video
Yusuf Aytar, Carl Vondrick, Antonio Torralba
cs.CVcs.LGcs.SDarXiv:1610.09001v12016MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel +3
cs.CVarXiv:2409.19152v12024Bit-pragmatic Deep Neural Network Computing
J. Albericio, P. Judd, A. Delmás +2
cs.LGcs.AIcs.ARarXiv:1610.06920v12016Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Yan Shu, Zheng Liu, Peitian Zhang +5
cs.CVarXiv:2409.14485v42024Learning Low Dimensional Convolutional Neural Networks for High-Resolution Remote Sensing Image Retrieval
Weixun Zhou, Shawn Newsam, Congmin Li +1
cs.CVarXiv:1610.03023v22016Point Cloud Completion by Skip-attention Network with Hierarchical Folding
Xin Wen, Tianyang Li, Zhizhong Han +1
cs.CVarXiv:2005.03871v22020VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Yecheng Wu, Zhuoyang Zhang, Junyu Chen +9
cs.CVcs.LGarXiv:2409.04429v32024OmniRe: Omni Urban Scene Reconstruction
Ziyu Chen, Jiawei Yang, Jiahui Huang +9
cs.CVarXiv:2408.16760v22024We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Runqi Qiao, Qiuna Tan, Guanting Dong +15
cs.AIcs.CLcs.CVarXiv:2407.01284v12024Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
Kailu Wu, Fangfu Liu, Zhihan Cai +5
cs.CVcs.GRcs.LGarXiv:2405.20343v32024AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels
Javier Tirado-Garín, Alan Savio Paul, Shuai Chen +5
cs.CVarXiv:2609.02798v12026Change Guiding Network: Incorporating Change Prior to Guide Change Detection in Remote Sensing Imagery
Chengxi Han, Chen Wu, Haonan Guo +3
cs.CVeess.IVarXiv:2404.09179v12024Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen, Yifei Huang, Jilan Xu +7
cs.CVarXiv:2403.09626v12024Frequency-Adaptive Dilated Convolution for Semantic Segmentation
Linwei Chen, Lin Gu, Ying Fu
cs.CVarXiv:2403.05369v72024SpatialGuard: Harness-Guided Verifiable Spatial Reasoning for Text-to-Image Generation
Ziyun Qian, Zizhi Chen, Yizhou Liu +3
cs.CVarXiv:2609.01582v12026Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
Liting Lin, Heng Fan, Zhipeng Zhang +3
cs.CVarXiv:2403.05231v22024Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann +14
cs.CVarXiv:2403.03206v12024PointMamba: A Simple State Space Model for Point Cloud Analysis
Dingkang Liang, Xin Zhou, Wei Xu +5
cs.CVarXiv:2402.10739v52024HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin +9
cs.LGcs.AIcs.CLarXiv:2402.04249v22024ATGV-Net: Accurate Depth Super-Resolution
Gernot Riegler, Matthias Rüther, Horst Bischof
cs.CVarXiv:1607.07988v12016VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality
Ying Jiang, Chang Yu, Tianyi Xie +8
cs.HCcs.CVarXiv:2401.16663v22024RGBD Salient Object Detection via Deep Fusion
Liangqiong Qu, Shengfeng He, Jiawei Zhang +3
cs.CVarXiv:1607.03333v12016DriveLM: Driving with Graph Visual Question Answering
Chonghao Sima, Katrin Renz, Kashyap Chitta +7
cs.CVarXiv:2312.14150v32023Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
Huan Ling, Seung Wook Kim, Antonio Torralba +2
cs.CVcs.LGarXiv:2312.13763v22023Putting the Object Back into Video Object Segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price +2
cs.CVarXiv:2310.12982v22023HIC-YOLOv5: Improved YOLOv5 For Small Object Detection
Shiyi Tang, Shu Zhang, Yini Fang
cs.CVarXiv:2309.16393v22023ICAFusion: Iterative Cross-Attention Guided Feature Fusion for Multispectral Object Detection
Jifeng Shen, Yifei Chen, Yue Liu +3
cs.CVarXiv:2308.07504v12023OCR-Based Field Extraction for Archaeological Pottery Metadata: The CENTURIA Dataset
Gissu Valentina Naghavi, Dominik Hagmann, Martin Kampel +1
cs.CVarXiv:2608.30616v12026Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols
Tiffanie Godelaine, Maxime Zanella, Karim El Khoury +2
cs.CVcs.AIarXiv:2608.30420v12026Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving
Tian Zhang, Zhuo Huang, Hongrui Ye +3
cs.CVcs.AIcs.LGarXiv:2608.30122v12026RGBD Datasets: Past, Present and Future
Michael Firman
cs.CVcs.ROarXiv:1604.00999v22016A Survey of Techniques for Optimizing Transformer Inference
Krishna Teja Chitty-Venkata, Sparsh Mittal, Murali Emani +2
cs.LGcs.ARcs.CLarXiv:2307.07982v12023Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
Yu. A. Malkov, D. A. Yashunin
cs.DScs.CVcs.IRarXiv:1603.09320v42016Generating images with recurrent adversarial networks
Daniel Jiwoong Im, Chris Dongjoo Kim, Hui Jiang +1
cs.LGcs.CVarXiv:1602.05110v52016Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation
Mukul Khanna, Yongsen Mao, Hanxiao Jiang +7
cs.CVarXiv:2306.11290v32023Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
Xingang Pan, Ayush Tewari, Thomas Leimkühler +3
cs.CVcs.GRarXiv:2305.10973v22023MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
Mingdeng Cao, Xintao Wang, Zhongang Qi +3
cs.CVarXiv:2304.08465v12023GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
Guangting Zheng, Yiyuan Zhang, Tao Yang +4
cs.CVarXiv:2608.29335v12026SGRNet: Spatially Guided Radiology Network for Structured Radiological Reporting of Head and Neck Cancer
Ayush Gupta, Vinkle Srivastav, Prateek Upadhya +3
cs.CVarXiv:2608.29153v12026Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks
Lin Wang, Kuk-Jin Yoon
cs.CVcs.AIcs.LGarXiv:2004.05937v72020STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation
Shingeon Kim, Hyeyoon Lee, Dain Kwon +5
cs.CVcs.AIarXiv:2608.29145v12026Cross-dimensional Weighting for Aggregated Deep Convolutional Features
Yannis Kalantidis, Clayton Mellina, Simon Osindero
cs.CVarXiv:1512.04065v22015Projection-Aware End-to-End Learned Video Compression for 360-Degree Video
Niloofar Maani
cs.CVarXiv:2608.28689v22026FreeNeRF: Improving Few-shot Neural Rendering with Free Frequency Regularization
Jiawei Yang, Marco Pavone, Yue Wang
cs.CVarXiv:2303.07418v12023Stabilizing Transformer Training by Preventing Attention Entropy Collapse
Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin +5
cs.LGcs.AIcs.CLarXiv:2303.06296v22023Exploiting Deep Generative Prior for Versatile Image Restoration and Manipulation
Xingang Pan, Xiaohang Zhan, Bo Dai +3
eess.IVcs.CVarXiv:2003.13659v42020Dynamic Multiscale Graph Neural Networks for 3D Skeleton-Based Human Motion Prediction
Maosen Li, Siheng Chen, Yangheng Zhao +3
cs.CVcs.LGstat.MLarXiv:2003.08802v12020Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
Jiazhao Liang, Hao Huang, Shuaihang Yuan +8
cs.CVarXiv:2608.28218v12026Rapid AI Development Cycle for the Coronavirus (COVID-19) Pandemic: Initial Results for Automated Detection & Patient Monitoring using Deep Learning CT Image Analysis
Ophir Gozes, Maayan Frid-Adar, Hayit Greenspan +5
eess.IVcs.CVcs.LGarXiv:2003.05037v32020A Geometry-Driven, Framework-Agnostic Optimization for Object Pose Estimation
Wei Chen, Tao Zhen, Zhongchen Shi +3
cs.CVarXiv:2608.26859v12026Calibrating Deep Neural Networks using Focal Loss
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal +3
cs.LGcs.CVstat.MLarXiv:2002.09437v22020EcoTTA: Memory-Efficient Continual Test-time Adaptation via Self-distilled Regularization
Junha Song, Jungsoo Lee, In So Kweon +1
cs.CVarXiv:2303.01904v42023Spatially-Adaptive Feature Modulation for Efficient Image Super-Resolution
Long Sun, Jiangxin Dong, Jinhui Tang +1
cs.CVarXiv:2302.13800v12023Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation
Jian Liang, Dapeng Hu, Jiashi Feng
cs.CVcs.LGarXiv:2002.08546v62020Canadian Adverse Driving Conditions Dataset
Matthew Pitropov, Danson Garcia, Jason Rebello +4
cs.CVarXiv:2001.10117v32020