# Africa Language AI Evaluator
A research-driven evaluation framework for assessing AI model performance on African languages, with a focus on Kiswahili and Sheng.
## Live Demo
your-app.vercel.app
API Documentation:
your-api.onrender.com
Note: When external model APIs are unavailable, the system uses simulated candidates to demonstrate evaluation, comparison, and failure analysis workflows.
---
## Positioning
This system functions as an evaluation infrastructure layer for African language model analysis. It does not only score outputs, but identifies failure patterns and supports model and dataset improvement strategies.
---
## Problem
Most AI language models are evaluated using surface-level metrics such as BLEU, which fail to capture:
- Cultural meaning
- Informal speech patterns
- Code-switching behavior (e.g. Sheng)
- Contextual correctness
This results in poor performance in real-world African language use cases.
---
## Capabilities
| Feature | Description |
|--------|------------|
| Auto Evaluation Lab | Generates and evaluates multiple candidate outputs |
| Model Comparison | Ranks outputs using a structured scoring framework |
| Failure Analysis | Identifies meaning loss, tone mismatch, and cultural errors |
| Research Insights | Aggregates evaluation trends and patterns |
| Evaluation History | Stores and filters past evaluations |
| Dataset Import | Supports CSV-based batch evaluation |
---
## Evaluation Dimensions
| Dimension | Scale | Description |
|----------|------|------------|
| Meaning Preservation | 1–5 | Semantic fidelity between source and output |
| Fluency | 1–5 | Natural readability in the target language |
| Cultural Context | 1–5 | Cultural and linguistic appropriateness |
| Safety & Bias Risk | low / medium / high | Detection of harmful or biased framing |
---
## Supported Languages
| Language | Notes |
|---------|------|
| English | Source and target |
| Kiswahili | Formal tone and politeness handl …