Source-linked AI summary

Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency

Sakib Shahriar, Brady Lund, Nishith Reddy Mannuru, Muhammad Arbab Arshad, Kadhim Hayawi, Ravi Varma Kumar Bevara, Aashrith Mannuru, Laiba Batool

arXiv:2407.09519v1cs.AIcs.CL

TL;DR

The paper addresses limited comprehensive evidence about GPT-4o’s capabilities across language, vision, speech, and multimodal tasks. It evaluates the model using exams, reasoning and translation tasks, image and audio assessments, and multimodal benchmarks, finding strong but variable performance alongside limitations with complex and ambiguous inputs. The authors also identify dataset breadth and the absence of human judgment as important evaluation constraints.

  • Problem

    Comprehensive evaluation of GPT-4o remained limited across its language, vision, speech, and multimodal capabilities.

  • Method

    The study evaluates GPT-4o with language assessments, semantic-similarity analysis, image and audio tasks, and visual question answering using paired images and questions.

  • Results

    GPT-4o shows strong performance across multiple evaluated domains, while its multimodal assessment examines integration of visual and linguistic information through tasks such as visual question answering.

  • Takeaways & Limitations

    The findings support evaluating GPT-4o beyond language and reasoning because its newer vision, speech, and cross-modal capabilities create additional assessment dimensions.

  • Takeaways & Limitations

    The evaluation uses relatively small, non-exhaustive image and audio datasets, omits image and audio generation, lacks human judgment, and finds inconsistencies on complex or ambiguous inputs.

Abstract

from arXiv · show

As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study comprehensively evaluates the language, vision, speech, and multimodal capabilities of GPT-4o. The study employs standardized exam questions, reasoning tasks, and translation assessments to assess the model's language capability. Additionally, GPT-4o's vision and speech capabilities are tested through image classification and object recognition tasks, as well as accent classification. The multimodal evaluation assesses the model's performance in integrating visual and linguistic data. Our findings reveal that GPT-4o demonstrates high accuracy and efficiency across multiple domains in language and reasoning capabilities, excelling in tasks that require few-shot learning. GPT-4o also provides notable improvements in multimodal tasks compared to its predecessors. However, the model shows variability and faces limitations in handling complex and ambiguous inputs, particularly in audio and vision capabilities. This paper highlights the need for more comprehensive benchmarks and robust evaluation frameworks, encompassing qualitative assessments involving human judgment as well as error analysis. Future work should focus on expanding datasets, investigating prompt-based assessment, and enhancing few-shot learning techniques to test the model's practical applicability and performance in real-world scenarios.

1. Introduction

GPT-4o emerged amid rapid progress in large language models, but its broad capabilities had not been comprehensively evaluated. This study addresses that gap across language, vision, speech, and multimodal tasks.

  • Research purpose: GPT-4o is evaluated across language, vision, speech, and multimodal capabilities to understand its strengths and limitations.The study also considers performance relative to previous GPT models and contemporary systems.
  • Context: GPT-4o is one large language model within the broader hierarchy of artificial intelligence, machine learning, deep learning, generative AI, and LLMs.The paper situates GPT-4o as the latest version of OpenAI’s generative pre-trained transformer model.
  • Motivation: The study motivates careful evaluation because LLMs can produce inaccurate, biased, harmful, or misleading information and may contribute to misinformation.The paper connects close imitation of human language with both practical applications and risks requiring scrutiny.
  • Research gap: Evaluation of GPT-4o remained limited despite studies of threats, diagnosis, multilingual ability, public sentiment, and related chatbot applications.One cited study found GPT-4o underperformed Claude 3 Opus in radiology diagnosis.
  • Expanded scope: GPT-4o’s vision, speech, and cross-modal capabilities enable evaluation beyond the language and reasoning tasks used for earlier GPT versions.GPT-4V had previously been tested on vision tasks, while speech was new to GPT-4o.

2. Language Capacity of GPT-4o

The language evaluation examines GPT-4o’s ability to understand and generate language through standardized exams, reasoning tasks, and translation activities.

  • Evaluation scope: GPT-4o’s language capacity is assessed through exams, reasoning tasks, and translation activities.These tasks target different aspects of processing and producing coherent, contextually appropriate language.
  • Exam evaluation: Standardized and board exam questions test GPT-4o’s ability to comprehend complex problems and produce coherent, relevant, accurate responses.The questions span multiple subjects and measure proficiency with structured tasks.
  • Exam evaluation: The evaluation analyzes GPT-4o’s generated responses to questions drawn from varied standardized and board examinations.The approach is designed to measure structured question handling across subjects.

Performance on USMLE

On the USMLE Step 1 evaluation, GPT-4o answered 98 of 118 questions correctly, achieving 83.1% accuracy. Its performance exceeded GPT-3.5 but was below GPT-4.

  • Results: 83.1% accuracy came from 98 correct answers out of 118 USMLE questions.The result is reported as GPT-4o’s performance on the evaluation.
  • Model comparison: GPT-4o outperformed GPT-3.5’s 51.67% accuracy but remained below GPT-4’s 90.00%.The comparison is reported against predecessor-model results.
  • Interpretation: The paper attributes GPT-4o’s lower result relative to GPT-4 to a design focus on efficiency and speed, while GPT-4 targets more complex tasks.This interpretation is stated in the reported discussion of the USMLE results.
  • Practical implication: The authors describe GPT-4o as suitable for medical education requiring fast, interactive feedback, while GPT-4 handles more intricate questions.The proposed educational use is tied to GPT-4o’s reported accuracy and efficiency.

Performance on CFA

On a CFA Level 1 mock exam, GPT-4o answered 76 of 89 questions correctly, yielding 85.39% accuracy. The evaluation used zero-shot prompting without hints or specific instructions.

  • Assessment scope: The CFA Level 1 exam covers foundational finance and investment knowledge, including ethics, quantitative methods, economics, and portfolio management.It tests both theoretical understanding and practical application.
  • Results: GPT-4o correctly answered 76 of 89 CFA Level 1 mock-exam questions, achieving 85.39% accuracy.The dataset was designed to mirror the style and difficulty of the actual CFA Level 1 exam.
  • Evaluation setup: The CFA evaluation used zero-shot prompting without hints or specific instructions.The comparison was conducted under the same zero-shot framing described in the passage.

Performance on SAT

GPT-4o was evaluated on SAT Practice Test #1 across reading, writing, and math modules, with comparison to earlier GPT models. It achieved the highest reported Reading & Writing accuracy and strong Math accuracy, though GPT-4 remained slightly higher in Math.

  • GPT-4o was evaluated using SAT Practice Test #1, which included reading, writing, and math questions across two modules.
  • 90.91% was GPT-4o's highest accuracy in Reading & Writing, surpassing all older models.
  • 87.04% was GPT-4o's accuracy in Math, slightly below GPT-4 but above the other compared models.
  • Figures 2–5 illustrated both correct and incorrect GPT-4o responses across the SAT categories.

Performance on MBE

The supplied passages combine MBE performance with the paper's broader reasoning-evaluation context. GPT-4o performed comparably to GPT-4 on the MBE and showed strong results across deductive, inductive, and abductive reasoning, while response consistency remained a concern.

  • Performance on MBE: GPT-4o answered 15 of 20 MBE questions correctly, achieving 75% accuracy.
  • Performance on MBE: GPT-4o performed comparably to GPT-4 on the MBE and substantially exceeded GPT-3.5's reported 45.10% accuracy.
  • Reasoning: The reasoning evaluation covered deductive, inductive, and abductive tasks using five datasets.
  • Reasoning: GPT-4o achieved nearly flawless deductive results, perfect bAbI task 16 performance, 17/30 on CLUTRR, and 27/30 on αNLI.
  • Limitations: GPT-4o sometimes produced different answers across repeated sessions or requested users to choose among answers, indicating difficulty with ambiguity and consistency.
  • Limitations: The paper suggests that further study should examine broader and more intricate domain-specific reasoning problems.

Data

The translation evaluation drew from established multilingual datasets and used random samples to assess GPT-4o across six languages. The sampling design aimed to balance feasibility with linguistic and challenge diversity.

  • Spanish, Arabic, French, Portuguese, and Russian data came from OPUS, while Hindi data came from the IIT Bombay English-Hindi Parallel Corpus.
  • 500 data points were randomly sampled from each dataset to balance feasibility with coverage of sentence structures, vocabulary, and translation challenges.

Evaluation Method

The evaluation measures semantic similarity between reference and GPT-4o translations by embedding sentences with BERT-based models and comparing their vectors using cosine similarity.

  • Semantic Similarity: Sentence embeddings from paraphrase-MiniLM-L6-v2 represent the semantic information of reference and generated translations for comparison.The model is a pre-trained sentence-transformers model based on BERT.
  • Cosine Similarity: Cosine similarity measures sentence similarity through the angle between their embedding vectors.Vectors pointing in the same direction indicate greater semantic similarity, while divergent directions indicate less similarity.
  • Cosine Similarity: Cosine similarity values range from -1 to 1, with 1 indicating identical vectors, 0 orthogonality, and -1 opposing vectors.These values provide an interpretable scale for comparing sentence embeddings.
  • Cosine Similarity: The cosine-similarity formula divides the dot product of embeddings A and B by the product of their magnitudes.A and B denote the embeddings of the two sentences.

Results

GPT-4o produced generally strong translation results across six languages, with the highest scores in Spanish and Portuguese and lower scores in Arabic and French. The evaluation also notes sampling and metric limitations that constrain interpretation.

  • Translation Results: 88% Spanish and 86% Portuguese translation accuracy were the highest results among the six evaluated languages.The other reported scores were 82% for Hindi, 80% for Russian, 78% for Arabic, and 75% for French.
  • Translation Results: Arabic and French achieved lower translation accuracy scores of 78% and 75%, respectively.The passage attributes these challenges to complex linguistic structures and nuances.
  • Interpretation: GPT-4o’s translation quality approached that of dedicated translation systems despite not being specifically optimized for translation.The comparison is based on similarity scores against existing translations.
  • Limitations: Sampling 500 data points per dataset may not fully represent each language’s linguistic diversity and complexity.Different random samples could produce different results, motivating larger and more representative datasets.
  • Limitations: BERT-based embeddings and cosine similarity may not capture cultural and contextual subtleties of translation quality.The paper recommends larger datasets, more language pairs, and advanced fine-tuning for future work.

3. Vision Capacity of GPT-4o

The vision evaluation tests GPT-4o on image classification, medical-image recognition, captioning, and few-shot tasks using curated image subsets, standard metrics, and qualitative error analysis. Performance is strong on fruit classification but varies substantially across domains, with major weaknesses in cancer detection and longer captions.

  • Evaluation Setup: Approximately 100 representative images were used per vision task, with prompts specifying the desired output and few-shot examples provided for selected tasks.Outputs were compared with ground-truth labels using metrics such as accuracy, alongside qualitative analysis of strengths and failure modes.
  • Fruit Classification: GPT-4o achieved average precision, recall, and F1-score of 0.98 on 10-class fruit classification.Papaya, Apple, Litchi, Hog Plum, Grapes, and Guava each received perfect precision, recall, and F1-scores of 1.0.
  • Drowsiness Classification: GPT-4o achieved average precision, recall, and F1-score of 0.80 on the drowsy-versus-natural image task without fine-tuning.This was lower than specialized VGG, ResNet, and CNN models but notable without domain-specific training.
  • Crop Disease Classification: Crop-disease classification produced average precision of 0.77, recall of 0.71, and F1-score of 0.68 across 20 classes.The paper attributes weaker performance in some classes to the need for specialized training.
  • Few-Shot Learning: Few-shot examples substantially improved normal-class recognition initially, but additional examples produced diminishing returns after one shot.The glaucoma class maintained a relatively high F1-score across shot levels with only slight improvement.
  • Medical Image Classification: Cancer, tumor, and aneurysm detection produced average precision of 0.21, recall of 0.32, and F1-score of 0.26.The model completely failed to predict the cancer class and frequently confused aneurysm with tumor.
  • Image Captioning: Image captioning achieved BLEU-1 of 0.193, declining to BLEU-4 of 0.031 as n-gram length increased.The decline indicates difficulty maintaining coherence and context in longer generated descriptions.

4. Speech Capacity of GPT-4o

The speech evaluation examines emotion detection in Arabic audio and accent classification across diverse English accents. Results are variable, with strong class confusion in emotion recognition and particularly frequent Malayalam–Telugu accent misclassification.

  • Emotion Detection: The Arabic emotion dataset contains 1,384 recordings labeled as happy, angry, or surprised.GPT-4o was evaluated on detecting discrete emotions in Arabic speech.
  • Emotion Detection: The model correctly predicted 21 surprised instances but only two angry instances, frequently confusing happy with surprised.The confusion matrix indicates substantial variation across emotion classes.
  • Accent Classification: AccentDB covers Bangla, Malayalam, Odiya, Telugu, metropolitan Indian, American, Australian, British, and Welsh English accents.The dataset provides varied linguistic backgrounds and phonetic and prosodic patterns for speech recognition evaluation.
  • Accent Classification: Malayalam was frequently misclassified as Telugu, while Bangla and Telugu also showed substantial confusion involving Malayalam.The paper connects these errors to potentially similar acoustic features and recommends more distinctive feature extraction and augmented training data.

5. Multimodal Capacity of GPT-4o

GPT-4o is evaluated on multimodal tasks that integrate visual and linguistic information, including VQA and vision-language benchmarks. Results indicate strong vision-language performance across diverse capabilities, alongside variability in few-shot VQA accuracy.

  • Multimodal evaluation: The multimodal evaluation examines how GPT-4o integrates information from sources such as text, images, and audio.This capacity is assessed because combining modalities can produce more comprehensive and contextually enriched responses.
  • Visual Question Answering: The VQA evaluation sampled 100 image-question pairs requiring GPT-4o to interpret images and answer natural-language questions.VQA combines computer vision and natural-language processing by pairing images with questions about their visual content.
  • Visual Question Answering: 0.36 accuracy was the peak VQA performance, achieved with eight shots, while performance decreased with one example.The reported pattern shows that adding a small number of examples did not improve performance uniformly across shot counts.
  • Visual Question Answering: The one-shot VQA decrease may reflect skewed answer distributions caused by an unrelated task in a setting with diverse answer possibilities.This explanation is presented as a possible reason for the observed variability.
  • Vision-language evaluation: MM-Vet evaluates integrated vision-language capabilities including recognition, OCR, knowledge, language generation, spatial awareness, and mathematics.Its tasks can require recognizing objects, reading image text, interpreting spatial relationships, using knowledge, and generating coherent responses.
  • Vision-language evaluation: GPT-4o outperforms previous models across all evaluated MM-Vet metrics, with particularly high performance in knowledge, spatial awareness, and language generation.The benchmark results are presented as evidence of advances in GPT-4o's vision-language capabilities compared with its predecessors.

6. Implications, Limitations, and Future Work

The study identifies practical promise for GPT-4o while emphasizing that current evaluations remain incomplete and that the model has weaknesses with complex or ambiguous inputs. Future work should broaden datasets and benchmarks, incorporate real-time and human-centered assessment, and refine few-shot evaluation.

  • Implications: GPT-4o's performance in medical question answering and financial analysis suggests utility for educational and professional training environments.Its vision-language integration may also support multimodal analysis in healthcare, finance, and customer service.
  • Implications: Comprehensive evaluation remains necessary because existing assessments do not fully capture advanced models' capabilities and weaknesses across real-world scenarios.The paper calls for broader benchmarks spanning more diverse data and tasks.
  • Limitations: The image and audio evaluation datasets were relatively small and non-exhaustive, limiting coverage of possible scenarios.The study assessed breadth across data types and multimodal inputs, but not depth within every category; image and audio generation were outside scope.
  • Limitations: The study did not use qualitative or human judgment, and GPT-4o showed inconsistent accuracy on ambiguous or complex inputs.The paper notes that quantitative metrics alone may miss practical usability and contextual accuracy.
  • Future Work: Real-time and longitudinal evaluations could clarify GPT-4o's adaptability, stability, and practical reliability in dynamic settings.Examples include monitoring driver drowsiness and detecting sudden changes in patient health through medical imaging.
  • Future Work: Future research should expand datasets, develop comprehensive benchmarks, explore additional multimodal inputs, and refine few-shot learning techniques.These directions are intended to improve understanding of model performance and support more reliable AI systems.
Loading 2407.09519v1…