Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

Β© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

tsegaimerhawi/tigrinya-web-scraper

Domain:

natural language processing

Record type:

software
Creator:
tse
Host:
# πŸ“° Tigrinya News **Browse Haddas Ertra and ask questions in Tigrinya or Englishβ€”powered by RAG and the Ge'ez script.** A full pipeline for **Tigrinya news**: scrape PDFs from Haddas Ertra, extract and clean Ge'ez text, embed with **LlamaIndex** + **Gemini**, store in **Qdrant**, and query via a **React + FastAPI** app. Built for researchers, linguists, and anyone working with Tigrinya NLP and low-resource language tech. --- ## ✨ What You Can Do | **In the app** | **In Script Runner** | |----------------|----------------------| | πŸ“– Browse scraped articles with full text and metadata | πŸ•·οΈ Download Haddas Ertra PDFs (configurable limit) | | πŸ“‹ Copy text, run word frequency, stats, sentence extraction | πŸ“„ Extract Ge'ez text, NER, image descriptions β†’ `raw_data.json` | | πŸ€– **Ask questions in Tigrinya or English**β€”answers from the ingested corpus (RAG) | πŸ“¦ Ingest into Qdrant (LlamaIndex + Gemini embeddings) | | πŸ” Semantic search over the news corpus | βœ… Check Qdrant, validate metadata and raw data | The **app** (React + FastAPI) is for **reading and asking**. The **Script Runner** (or CLI) is for **scraping, processing, and ingesting**. Run the pipeline once (or periodically), then use the app to explore and query. --- ## πŸ—οΈ How It Fits Together ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Pipeline UI: localhost (or script_runner β”‚ β”‚ on :8765) β†’ Scraper, PDF Processor, Llama Ingest β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό pdf_metadata.json raw_data.json Qdrant β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ App: localhost (backend :8000) β”‚ β”‚ Articles (browse, copy, NLP tools) + Ask (RAG) β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` --- ## πŸ› οΈ Tech Stack - **Frontend:** React, TypeScript, Vite - **Backend:** FastAPI, Python 3.8+ - **Scraping:** Playwri …

Visit

github.com

Languages

Tigrigna