Welcome¶
This is the documentation of the data processing backbone of the Research Software Observatory, supporting large-scale monitoring of software FAIRness in the life sciences.
The pipeline consolidates and harmonizes metadata from multiple registries and repositories, enriches it with external information, and pre-computes the FAIRsoft indicators and other metrics displayed in the Software Observatory interface.
At a glance
Language: Python ≥ 3.9 (tested on 3.10; deployment image uses 3.12)
Execution: CLI (rsetl), or a Docker image for VM deployment
Dependencies: pydantic, tenacity, pymongo, ... (see more)
Database: MongoDB
Main stages: Transformation → Integration & disambiguation → Merge → Evaluation (incremental by default)
Enrichment sub-pipelines: SPDX · Publications · Service availability · Similarity
Maintained by: Spanish National Bioinformatics Institute
Quickstart¶
Clone and install:
git clone https://github.com/inab/research-software-etl.git
cd research-software-etl
pip install -e .
Each execution can run as a single stage or as part of the full workflow through the unified CLI command rsetl:
rsetl run
Use rsetl --help or go to the CLI docs for more information.
Next steps¶
- Installation & Configuration – Set up the environment and dependencies.
- Main Pipeline Stages – Detailed description of each processing step.
- CLI reference - Learn how to run the pipeline
- Deployment – Run the pipeline as a container on a VM.
- Development Guide – Learn the project’s structure and how to extend it.
Next step → Installation & Configuration