Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AI on the Frontline: Evaluating Large Language Models in Real-World Conflict Resolution

Domain:

peace and securitynatural language processing

Record type:

paper
Creator:
Institute for Integrated Transitions
Editor:
Freeman, MarkBus
Publisher:
Zenodo
Host:avatar
This groundbreaking study authored by Nathalie Bussemaker and Mark Freeman and published by the Institute for Integrated Transitions (IFIT) reveals that all major large language models (LLMs) are providing dangerous conflict resolution advice without conducting basic due diligence that any human mediator would consider essential. IFIT tested six leading AI models including ChatGPT, Deepseek, Grok, and others on three real-world prompt scenarios from Syria, Sudan, and Mexico. Each LLM response, generated on June 26, 2025, was evaluated by two independent five-person teams of IFIT researchers across ten key dimensions, based on well-established conflict resolution principles such as due diligence and risk disclosure. Scores were assigned on a 0 to 10 scale for each dimension to assess the quality of each LLM’s advice.  A senior expert sounding board of IFIT conflict resolution experts from Afghanistan, Colombia, Mexico, Northern Ireland, Sudan, Syria, the United States, Uganda, Venezuela, and Zimbabwe then reviewed the findings to assess implications for real-world practice. From a total possible point value of 100/100, the average score across all six models was only 27 points. The maximum score was obtained by Google Gemini with 37.8/100, followed by Grok with 32.1/100, ChatGPT with 24.8/100, Mistral with 23.3/100, Claude with 22.3/100, and DeepSeek last with 20.7/100. All scores represent a failure to abide by minimal professional conflict resolution standards and best practices.  

Visit

doi.orgzenodo.org

Tags

AILLMlarge language modelconflict resolutionIFITpeacebuildingGoogle GeminiChatGPTDeepSeekMistral+10

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Large language models for frontline healthcare support in low-resource settingsMultimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative (Preprint)Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language BenchmarksEvaluating Metalinguistic Knowledge in Large Language Models across the World's LanguagesEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str

Multimodal large language models versus physicians for real-world inpatient diagnosis: Retrospective comparative (Preprint)

BACKGROUND Large language models (LLMs) are increasingly proposed for diagnostic

Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language Benchmarks

Democratization of AI is an important topic within the broader topic of the digital divide. This iss

Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages

LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structur

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (L