Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Malaria-Instruct: A Low-Resource Benchmark for Instruction-Tuned LLMs in Antimalarial Drug Discovery

Domaine:

healthcare

Type de record:

dataset
Créateur:
Ajala, MarvellousAdeAsh
Éditeur:
Zenodo
Hôte:avatar
Overview: This repository contains Malaria-Instruct, a specialized dataset designed to evaluate and fine-tune Large Language Models (LLMs) for molecular property prediction, specifically targeting antimalarial activity. While derived from ChEMBL, this dataset has been significantly restructured, filtered, and formatted for instruction-tuning and virtual screening tasks in resource-constrained research environments. Key Features: Instruction-Response Format: Data is structured as natural language prompts (SMILES strings as input, binary activity as output) to align with the training objectives of modern LLMs (e.g., Gemma, LlaSMO, Mistral). Rigorous Dissimilarity Splitting: Unlike standard random or scaffold splitting, this dataset employs a "Hard Split" based on molecular dissimilarity (Tanimoto coefficient <0.4 for train-test and 0.5 for train-val). This ensures that models are tested on their ability to generalize to novel chemical spaces, simulating real-world lead optimization. Neglected Disease Focus: Specifically curated from assays involving various strains of Plasmodium falciparum, addressing the data gap in AI for the Global South. It contains assay on different strains on Plasmodium falciparum Benchmarking Metadata: Includes pre-calculated baseline results using Morgan Fingerprints (2048-bit) and classical models (Random Forest, XGBoost). File Structure: MalariaData_bioactivity_LLM_fewshot_dataset.zip: A zip file containing few shots (1-5 shorts) datasets in csv file with the splits. MalariaData_bioactivity_selected_for_LLM_with_splits_fewshot.zip: A zip file containing the csv file of the dataset repeated splitted 5 times. classical_results.csv: pre-computed classical ML baselines result summary. README.md: Detailed documentation on filtering criteria and prompt templates. Use Case: This dataset is intended for researchers developing chemistry-aware LLMs, MLOps engineers optimizing small-scale models for low-compute environments, and medicinal chemists working on malaria drug discovery.

Visit

doi.org

Tags

Drug DiscoveryMalariaLarge Language ModelsInstruction tuningGlobal SouthCheminformaticsVirtual Screening

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Performance variation in instruction-tuned models on XTREME benchmark with low-resource intermediate-task trainingFew-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMsPrompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource LanguagesArtificial intelligence for antiviral drug discovery in low resourced settings : a perspectiveLLM Probe: Evaluating LLMs for Low-Resource Languageschinmayjainnnn/LLMs-for-Translation-of-Low-Resource-Languages

Performance variation in instruction-tuned models on XTREME benchmark with low-resource intermediate-task training

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs

This paper presents two effective approaches for Extractive Question Answering (QA) on the Quran. It

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

LLMs are typically trained in high-resource languages, and tasks in lower-resourced languages tend t

Artificial intelligence for antiviral drug discovery in low resourced settings : a perspective

Current antiviral drug discovery efforts face many challenges, including development of new drugs du

LLM Probe: Evaluating LLMs for Low-Resource Languages

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource a

chinmayjainnnn/LLMs-for-Translation-of-Low-Resource-Languages

Machine translation from assamese to english and vice versa using state of the art LLM's # Hindi-En