Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Urban Safety Perception Through the Lens of Large Multimodal Models: A Persona-based Approach

Domain:

socioeconomic

Record type:

paper
Creator:
BenLepLuc
Host:avatar
Understanding how urban environments are perceived in terms of safety is crucial for urban planning and policymaking. Traditional methods like surveys are limited by high cost, required time, and scalability issues. To overcome these challenges, this study introduces Large Multimodal Models (LMMs), specifically Llava 1.6 7B, as a novel approach to assess safety perceptions of urban spaces using street-view images. In addition, the research investigated how this task is affected by different socio-demographic perspectives, simulated by the model through Persona-based prompts. Without additional fine-tuning, the model achieved an average F1-score of 59.21% in classifying urban scenarios as safe or unsafe, identifying three key drivers of perceived unsafety: isolation, physical decay, and urban infrastructural challenges. Moreover, incorporating Persona-based prompts revealed significant variations in safety perceptions across the socio-demographic groups of age, gender, and nationality. Elder and female Personas consistently perceive higher levels of unsafety than younger or male Personas. Similarly, nationality-specific differences were evident in the proportion of unsafe classifications ranging from 19.71% in Singapore to 40.15% in Botswana. Notably, the model's default configuration aligned most closely with a middle-aged, male Persona. These findings highlight the potential of LMMs as a scalable and cost-effective alternative to traditional methods for urban safety perceptions. While the sensitivity of these models to socio-demographic factors underscores the need for thoughtful deployment, their ability to provide nuanced perspectives makes them a promising tool for AI-driven urban planning.

Visit

arxiv.org

Tasks

computer visionimage classification

Tags

Computers and SocietyArtificial Intelligence

Similar

Large language models through the lens of ubuntu for health research in sub-Saharan AfricaLarge Multimodal Models for Low-Resource Languages: A SurveyMangaUB: A Manga Understanding Benchmark for Large Multimodal ModelsAdvancing gender equity in ICT4D through a lens of care: a case study of Ethiopia through a care-based evaluative approachM3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsEnhancing Sustainability of Complex Epidemiological Models through a Generic Multilevel Agent-based Approach

Large language models through the lens of ubuntu for health research in sub-Saharan Africa

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panel

Advancing gender equity in ICT4D through a lens of care: a case study of Ethiopia through a care-based evaluative approach

There is increasing momentum across the Global South to use digital technologies to drive economic d

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

Enhancing Sustainability of Complex Epidemiological Models through a Generic Multilevel Agent-based Approach

International audience The development of computational sciences has fostered major a