# nlp_Darija_probabilistic_n-gram_language_model
What the code does:
-- Uses only:
- `darija-wiki`
- `goud.ma`
- `Youtube`
-- Cleans the text:
- removes URLs
- removes usernames
- handles hashtags
- lowercases Latin-script text
- removes extra noise
- tokenizes Darija text
- Builds a trigram probabilistic language model
- Uses Laplace smoothing
-- Splits the data into:
- 80% training
- 10% validation
- 10% testing
-- Reports:
- number of files used
- number of sentences
- number of tokens
- vocabulary size
- validation perplexity
- test perplexity
Generates sample Darija-like sentences from prompts such as:
--text:
-ana bghit
-واش
-hadchi