Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

sbkvnbs/NLP_biomed_ner_swahili

Domaine:

natural language processing

Type de record:

project
Créateur:
sbk
Hôte:
NLP project-Biomedical ner for Swahili # NLP_biomed_ner_swahili NLP project-Biomedical ner for Swahili Named Entity Recognition (NER) Model with Transformers Overview This project implements a Named Entity Recognition (NER) model using the Hugging Face Transformers library. The model is trained on tokenized text data and evaluated on a test dataset. It leverages pre-trained transformer models for token classification and uses the Weights & Biases (W&B) tool for experiment tracking. Requirements Ensure you have the following dependencies installed: pip install torch transformers datasets scikit-learn wandb Dataset Preparation The dataset is loaded as a DatasetDict from Pandas DataFrames: • train_df → Training dataset • val_df → Validation dataset • test_df → Test dataset Each dataset is converted into a Hugging Face Dataset format. Tokenization The dataset is tokenized using a tokenizer function that: • Tokenizes the Name column. • Aligns labels (NER_Category) with tokenized inputs. • Ensures padding and truncation up to 128 tokens. dataset = dataset.map(tokenize_function, batched=True, remove_columns=["Name", "NER_Category"]) Model Training A pre-trained model is loaded for token classification: model = AutoModelForTokenClassification.from_pretrained(model_name, num_labels=num_labels) Training parameters are set using TrainingArguments, specifying: • Learning rate, batch sizes, weight decay, and training epochs. • Mixed precision training (fp16=True). • W&B integration for logging. The model is trained using the Trainer class: trainer = Trainer( model=model, args=training_args, train_dataset=dataset['train'], eval_dataset=dataset['validation'], tokenizer=tokenizer ) Training is initiated using: trainer.train() Model Evaluation After training, the model is evaluated on the validation set: trainer.evaluate() To evaluate the model on the test dataset: • The test set is tokenized and converted to tensors. • Predictions and ground truth labels are extracted. • Labels are mapped back to NER tags, a …

Visit

github.com

Tasks

information extractionnamed entity recognition

Languages

Swahili