Logo Lanfrica

bpattath/llm-health-safety-evals

Domain:

healthcarenatural language processing

Record type:

software
Creator:
bpa
Host:
Open evaluation suite for LLM decision-support in low-resource health systems: grounding, multilingual safety (Malayalam), epidemic-signal reliability # LLM Health Safety Evals **An open evaluation suite for LLM decision-support in low-resource health systems: grounding, multilingual safety, and epidemic-signal reliability.** **Status: design phase complete, 12-item pilot harness runnable; seeking seed funding for the full build.** This repository contains the design document, task taxonomy with drafted example items, a dependency-free pilot runner (/pilot), and the roadmap. Code and datasets are developed here in the open. ## The problem Large language models are being wired into health decision-making in low- and middle-income countries, as chat interfaces over health data, triage aids, and advisory tools, faster than anyone is checking whether they are safe in those contexts. Three gaps stand out: 1. **Grounding.** When an assistant is constrained to a verified health-data store, does that grounding actually hold under realistic queries from health officials, and how often do models fabricate statistics when the data is missing? 2. **Multilingual safety.** Safety behaviours (refusals, caveats, calibration) are measurably weaker outside English. Health staff in Kerala, India (the suite's deployment context) work in Malayalam and English–Malayalam code-switching. No evaluation exists for whether guardrails survive the languages of actual deployment. 3. **Epidemic-signal reliability.** After the 2025 collapse of global surveillance funding, AI assistants are becoming the interpretive layer over incomplete, degraded surveillance data. Can models interpret potential outbreak signals honestly, expressing calibrated uncertainty and declining to conclude when the evidence is insufficient? This suite evaluates all three, and publishes everything (harness, task datasets, results across frontier and open-weight models) under open licences, so any team deploying LLMs into health systems can test before shipping. ## Deployment context The suite is designed against **First Signal**, a climate–health early-warning pla …

Licenses