Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,561 to 13,620 of 18,776
Generative Multimodal Models are In-Context Learners
Quan Sun, Yufeng Cui, Xiaosong Zhang +8
cs.CVarXiv:2312.13286v22023Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
Christopher Clark, Mark Yatskar, Luke Zettlemoyer
cs.CLcs.CVcs.LGarXiv:1909.03683v12019PoseTrack: A Benchmark for Human Pose Estimation and Tracking
Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov +4
cs.CVarXiv:1710.10000v22017InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
Jiale Xu, Weihao Cheng, Yiming Gao +3
cs.CVarXiv:2404.07191v22024DeeperGCN: All You Need to Train Deeper GCNs
Guohao Li, Chenxin Xiong, Ali Thabet +1
cs.LGcs.CVstat.MLarXiv:2006.07739v12020LangSplat: 3D Language Gaussian Splatting
Minghan Qin, Wanhua Li, Jiawei Zhou +2
cs.CVarXiv:2312.16084v22023PARE: Part Attention Regressor for 3D Human Body Estimation
Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges +1
cs.CVarXiv:2104.08527v22021Pixel Difference Networks for Efficient Edge Detection
Zhuo Su, Wenzhe Liu, Zitong Yu +5
cs.CVarXiv:2108.07009v12021DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification
Hongrun Zhang, Yanda Meng, Yitian Zhao +4
cs.CVcs.AIcs.LGarXiv:2203.12081v12022Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
Haoyuan Gao, Junhua Mao, Jie Zhou +3
cs.CVcs.CLcs.LGarXiv:1505.05612v32015Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
Denis Tome, Chris Russell, Lourdes Agapito
cs.CVarXiv:1701.00295v42017Face Aging With Conditional Generative Adversarial Networks
Grigory Antipov, Moez Baccouche, Jean-Luc Dugelay
cs.CVarXiv:1702.01983v22017See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks
Xiankai Lu, Wenguan Wang, Chao Ma +3
cs.CVarXiv:2001.06810v12020Online Continual Learning in Image Classification: An Empirical Survey
Zheda Mai, Ruiwen Li, Jihwan Jeong +3
cs.LGcs.CVarXiv:2101.10423v42021FOTS: Fast Oriented Text Spotting with a Unified Network
Xuebo Liu, Ding Liang, Shi Yan +3
cs.CVarXiv:1801.01671v22018Understanding Diffusion Models: A Unified Perspective
Calvin Luo
cs.LGcs.CVarXiv:2208.11970v12022DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision
Lu Ling, Yichen Sheng, Zhi Tu +17
cs.CVcs.AIarXiv:2312.16256v22023CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search
Xiaolong Sun, Qichao Wang, Hangyu Li +1
cs.AIcs.CVarXiv:2608.25023v12026Stochastic Separability of Embedding Manifolds
Liqing Zhang
cs.LGcs.CVarXiv:2608.22874v12026Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges
Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
cs.CVarXiv:2608.23845v12026Investigating Relational Reasoning in VLMs
Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan +1
cs.CVarXiv:2608.23518v12026MorphoCLIP: Text-Supervised Contrastive Learning for Perturbation Matching in Cell Painting Images
Sukhrobbek Ilyosbekov, Shubham Gajjar, Rongfei Jin
cs.CVarXiv:2608.22690v12026Targeting the Attention Heads Behind Object Hallucination in LLaVA
Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
cs.CVcs.AIarXiv:2608.24966v12026FlashNormal: Detailed Surface Normal Estimation from Flash and No-Flash Images
Ruiyang Chen, Feiran Li, Heng Guo +1
cs.CVarXiv:2608.25360v12026Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos
Jheng-Ling Lee, Shang-Tse Chen
cs.CVcs.LGarXiv:2608.24093v12026When Should a Network Emit Geometry, and When Should It Detect It? Readout, Reconciliation, and Representation in Floorplan Vectorization
He Zhang
cs.CVarXiv:2608.25608v12026PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction
Mary Raymond
cs.CVcs.AIarXiv:2608.25401v12026RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
Bojia Zi, Xiaoyan Yang, Yu Zhou +7
cs.CVarXiv:2608.26101v12026A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
Kaichen Li, Zhilin Zhu, Jianhao Huang +7
cs.CVcs.AIarXiv:2608.26095v12026DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors
Tuo Chen, Jie Gui, Minjing Dong +4
cs.CVarXiv:2608.25851v12026Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution
Ren Wang, Yung-Yu Chuang
cs.CVarXiv:2608.25998v120264DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting
Yueen Ma, Zenglin Xu, Irwin King
cs.CVarXiv:2608.25956v12026TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding
Yuqiang Lin, Yan Shi, Sam Lockyer +5
cs.CVcs.AIarXiv:2608.25935v12026UltraPIPS: Improving model perception in B-mode ultrasound with foundation models
Tal Grutman, Tali Ilovitsh
cs.CVeess.IVarXiv:2608.26033v12026Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18
cs.CVarXiv:2608.25924v12026Label-Free Foundational Model Selection for Medical Image Classification under Distribution Shift via Pseudo Label Discrepancy
Juan Iñaki Larrea, Lucas Mansilla, Enzo Ferrante
cs.CVarXiv:2608.25810v12026Embedding NDRE Trajectories into Contrastive Learning for Label-Free, Physiology-Aware Crop-Stress Staging and DSS Outputs
Shafqaat Ahmad
cs.CVarXiv:2608.25888v12026LUTSeg: A Longitudinal Multi-Expert Dataset for Ulcer Tissue Segmentation
Karen Sanchez, Carlos Hinojosa, Albert A. Ávila +5
cs.CVarXiv:2608.25866v12026Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Jiaming Zhou, Qihang Zhang, Gangwei Xu +9
cs.ROcs.CVarXiv:2608.26103v12026Auditable CT Phenotyping Through Report-derived Radiological Observations
Riga Wu, Walter Witschey, Yicheng Li +9
cs.CVarXiv:2608.25948v12026Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
Grace W. Lindsay
q-bio.NCcs.CVcs.NEarXiv:2001.07092v22020Learning Late, Guiding Early: Timestep-Decoupled Semantic Guidance for Fair Face Generation
Subir Kumar Parida, Rajbabu Velmurugan, Ketan Kotwal +2
cs.CVarXiv:2608.25862v12026Steer the Sampling, Not the Kernel Grid: Geometry-Guided Sampling Operator for Volumetric Segmentation
Sizhe Wang, Himashi Peiris, Zhaolin Chen
cs.CVarXiv:2608.25819v12026TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection
Qiangqiang Zhou, Jiacong Yu, Jiawei Xu +3
cs.CVarXiv:2608.25808v12026THA-Flow Generative Model: Prosthesis Geometry Prediction from Preoperative CT
Yiping Wang, Jie Li, Jingyu Shen +1
cs.CVarXiv:2608.25845v12026FRAME: separating sampling variation from representational cause in medical imaging fairness
Mahshad Lotfinia, Daniel Truhn, Andreas Maier +1
cs.CVcs.AIcs.LGarXiv:2608.25981v12026Multi-Oriented Text Detection with Fully Convolutional Networks
Zheng Zhang, Chengquan Zhang, Wei Shen +3
cs.CVarXiv:1604.04018v22016InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control
Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli +5
cs.CVarXiv:2608.25734v12026Deep Learning Segmentation of Diffusion-Weighted MRI Acute Ischaemic Stroke: A Pragmatic Evaluation Across Three Datasets
Atle Bjørnerud, Till Schellhorn, Thor H. Skattør +4
cs.CVarXiv:2608.25675v12026CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery
Yuanpei Liu, Zhenqi He, Jialu Tang +1
cs.CVarXiv:2608.25692v12026Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Fuxiao Liu, Kevin Lin, Linjie Li +3
cs.CVcs.AIcs.CEarXiv:2306.14565v42023Precipitation Downscaling Using Foundation Model-Conditioned Diffusion
Victor Nascimento Ribeiro, Jorge Guevara, Jorge Sebastian Moraga +7
cs.CVcs.LGphysics.ao-pharXiv:2608.25858v12026State of the Art on Neural Rendering
Ayush Tewari, Ohad Fried, Justus Thies +16
cs.CVcs.GRarXiv:2004.03805v12020MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
Hao Yin, Paritosh Parmar, Lijun Gu +6
cs.CVcs.AIcs.ETarXiv:2608.26094v12026AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
Yudong Guo, Keyu Chen, Sen Liang +3
cs.CVarXiv:2103.11078v32021MIMONet: Multi-scale Input and Multi-scale Output Network for Salient Object Detection
Zhaojian Yao, Wei Gao, Tiesong Zhao +2
cs.CVarXiv:2608.25733v12026LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding
Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +3
cs.CVarXiv:2608.25729v12026Difficulty-Aware Sample Allocation for Adaptive Data Augmentation in Semantic Segmentation
Olasimbo Ayodeji Arigbabu, Abimbola Ismail Arigbabu
cs.CVcs.AIarXiv:2608.25710v12026VITAL: VIsual Tracking via Adversarial Learning
Yibing Song, Chao Ma, Xiaohe Wu +6
cs.CVarXiv:1804.04273v12018Controlling for Omitted Variable Bias in Deep Neural Networks
Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter +1
stat.MEcs.CVcs.LGarXiv:2608.25930v12026