Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

kimx3966/weighted-x-entropy-asr

Domain:

natural language processing

Record type:

software
Creator:
kim
Host:
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition # Weighted Cross-Entropy for Low-Resource Languages in Multilingual Speech Recognition This repository contains code for the paper titled "Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition". The paper addresses the challenge of integrating low-resource languages into multilingual automatic speech recognition (ASR) systems. ## Code Overview 📁 The repository includes code for the data preprocessing, augmentation, training, and evaluation of the Whisper multilingual ASR model. It also provides scripts for fine-tuning the model with weighted cross-entropy and language-specific data augmentation. Additionally, the repository contains dataset manipulation scripts, model configuration files, and example usage scripts. ## Usage 🚀 1. **Clone this repository:** ```bash git clone github.com cd wce-low-resource-language-multilingual-asr ``` 2. **Install the required dependencies:** ```bash pip install -r requirements.txt ``` 3. **Just run the `.sh` file:** ```bash ./run.sh ``` ## Setup ⚙️ The code is prepared for multilingual training on the languages of the paper. Feel free to modify or add new languages. ### Basic Configuration The basic configuration of the training and data set can be set in the `run.sh` file: ```bash --base_model_dir "openai/whisper-small" \ --output_dir "your_location" \ --dataset_dir "your_dataset_location" \ --load_dataset_from_disk True \ --save_dataset_to_disk False \ --prune_well_datasets False \ --augment_gl_data False \ --max_input_length 30 \ --min_input_length 0 \ ``` ### Training Parameters Below are the training-related parameters that can be configured: ```bash --per_device_train_batch_size 16 \ --gradient_accumulation_steps 1 \ --learning_rate 1e-5 \ --weight_decay 0.01 \ --warmup_steps 800 \ --max_steps 8000 \ --gradient_checkpointing True \ --fp16 True \ --evaluation_strategy "steps" \ --per_device_eval_batch_size 8 \ …

Visit

github.com

Tasks

automatic speech recognitionspeech processing

Licenses

Apache-2.0

Similar

Entropy-Weighted Water Quality Index Assessment of Groundwater in Ibadan Metropolis, NigeriaTsallis entropy-KETEWStammeka/afaan-oromo-entropy-2026DanDrappaChapiti/chichewa-frequency-entropy-analysisEntropy estimation and entropy-based encoding of written Amharic language for efficient transmission in telecom networksOn the Entropy of Written Spanish

Entropy-Weighted Water Quality Index Assessment of Groundwater in Ibadan Metropolis, Nigeria

Abstract An entropy-weighted water quality index (EWQI) was used in this study to evaluate

Tsallis entropy-KETEWS

ZENODO DEPOSIT — METADATA AND DESCRIPTION
==================================

tammeka/afaan-oromo-entropy-2026

# Afaan Oromo Entropy 2026

DanDrappaChapiti/chichewa-frequency-entropy-analysis

Python tool for frequency and Shannon entropy analysis of the Chichewa language.

Entropy estimation and entropy-based encoding of written Amharic language for efficient transmission in telecom networks

On the Entropy of Written Spanish

This paper reports on results on the entropy of the Spanish language. They are based on an analysi