Abstract / Summary
Bioinformatics analysis of non-model organisms remains challenging because of the limited availability of specialized tools, as most R packages are optimized for well-annotated model species. This problem is exacerbated by genomic and proteomic data being scattered across multiple databases, each employing different identifiers based on varying reference annotations. Moreover, some of these platforms, such as MIIP, HitPredict, ApicoTFDB, and MPMP, often suffer from manual query interfaces, outdated identifiers, and inefficient data retrieval methods, complicating their use in large-scale bioinformatic analyses. We present plasmoRUtils , an R package that enables data retrieval across major apicomplexan resources such as OrthoMCL, VEuPathDB, and its associated databases. The package also offers methods to access static databases, such as MIIP and ApicoTFDB, through web scraping, assists in ID conversion tasks, and provides utilities for enrichment analysis and parasite stage estimation using bulk or single-cell RNA-seq references. By automating data retrieval and transformation within RStudio, plasmoRUtils eliminates the need for manual database queries, facilitating end-to-end workflow development without leaving the R environment. plasmoRUtils is freely available at https://rohit-satyam.github.io/plasmoRUtils/ with comprehensive documentation and runs on all major platforms.