<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://index.biohackrxiv.org//feed/by_tag/BioHackEU25.xml" rel="self" type="application/atom+xml" /><link href="https://index.biohackrxiv.org//" rel="alternate" type="text/html" /><updated>2026-08-10T19:26:33+00:00</updated><id>https://index.biohackrxiv.org//feed/by_tag/BioHackEU25.xml</id><title type="html">BioHackrXiv Preprints</title><subtitle>Preprints for BioHackathons</subtitle><author><name>GitHub User</name><email>your-email@domain.com</email></author><entry><title type="html">Improving package annotation in metabolomics and proteomics via robust, ontology-driven LLM integration</title><link href="https://index.biohackrxiv.org//2026/04/14/x5v6b.html" rel="alternate" type="text/html" title="Improving package annotation in metabolomics and proteomics via robust, ontology-driven LLM integration" /><published>2026-04-14T00:00:00+00:00</published><updated>2026-04-14T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/04/14/x5v6b</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/04/14/x5v6b.html"><![CDATA[<p>Identifying the most appropriate bioinformatics tool for a task remains challenging across multiple domains. Annotating tools with EDAM ontology
terms (e.g. topics, operations, input / output data and formats) can help, but manual annotation is labour-intensive, error-prone, and difficult
to scale, particularly given the high rate of first-time package developers in academic environments. At BioHackathon Europe 2025, our team
explored how Large Language Models (LLMs) can assist this process through the Model Context Protocol (MCP), an emerging open standard that
specifies how LLMs call external functions, using metabolomics as a domain use case. We developed an MCP-based workflow that grounds tool
descriptions in the EDAM ontology (Ison et al., 2013), improving reproducibility and semantic precision. Two core modules, entry-point
specification and semantic text segmentation, were completed during the hackathon, while additional mapping, validation, and reporting
functions were outlined for follow-up development. Benchmarking integrated with the BioChatter framework (Lobentanzer et al., 2025) demonstrated
that MCP-assisted models outperform unconstrained baselines on initial tests using metabolomics packages from bio.tools (Ison et al., 2019).
Ongoing work will expand benchmarking datasets, refine term-mapping logic, and extend the workflow to proteomics, supporting scalable,
ontology-driven annotation across the ELIXIR ecosystem.</p>]]></content><author><name>Sebastian Lobentanzer</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Identifying the most appropriate bioinformatics tool for a task remains challenging across multiple domains. Annotating tools with EDAM ontology terms (e.g. topics, operations, input / output data and formats) can help, but manual annotation is labour-intensive, error-prone, and difficult to scale, particularly given the high rate of first-time package developers in academic environments. At BioHackathon Europe 2025, our team explored how Large Language Models (LLMs) can assist this process through the Model Context Protocol (MCP), an emerging open standard that specifies how LLMs call external functions, using metabolomics as a domain use case. We developed an MCP-based workflow that grounds tool descriptions in the EDAM ontology (Ison et al., 2013), improving reproducibility and semantic precision. Two core modules, entry-point specification and semantic text segmentation, were completed during the hackathon, while additional mapping, validation, and reporting functions were outlined for follow-up development. Benchmarking integrated with the BioChatter framework (Lobentanzer et al., 2025) demonstrated that MCP-assisted models outperform unconstrained baselines on initial tests using metabolomics packages from bio.tools (Ison et al., 2019). Ongoing work will expand benchmarking datasets, refine term-mapping logic, and extend the workflow to proteomics, supporting scalable, ontology-driven annotation across the ELIXIR ecosystem.]]></summary></entry><entry><title type="html">Minimal information standardization of phenomic experimental data in animals</title><link href="https://index.biohackrxiv.org//2026/04/10/ncrkm.html" rel="alternate" type="text/html" title="Minimal information standardization of phenomic experimental data in animals" /><published>2026-04-10T00:00:00+00:00</published><updated>2026-04-10T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/04/10/ncrkm</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/04/10/ncrkm.html"><![CDATA[<p>The current landscape of animal phenomics is characterised by a substantial lack of standardisation, hindering data reuse, reproducibility, and interoperability across
studies, all of which are particularly important in light of the 3Rs principles for animal experiments (replace, reduce, refine). Within ELIXIR, the Domestic Animals
Genome and Phenome Focus Group emerged to establish standardised practices that enhance the quality and interoperability of animal research data. In this context, the
ISA model presents a robust, domain-agnostic framework well-established in the life sciences for describing experimental metadata. Notably, other scientific communities,
such as the ELIXIR Plant and Metabolomics Communities (MIAPPE, PhenoMeNal), have successfully leveraged the ISA model to improve the consistency and usability of their
metadata. Our project aims to develop a minimal information checklist tailored specifically for phenomics, facilitating the integration of diverse datasets, including
recirculation systems in agriculture, and fostering collaborative research efforts. We will focus on various goals.Identifying essential aspects of animal phenotyping,
informed by existing frameworks and community input. We aim to produce a concise and practical checklist that can be readily adopted by researchers, and promote a
culture of standardisation.Mapping the checklist to the ISA model ensures alignment with established standards, promotes interoperability and facilitates data reuse
while improving the overall quality of research outputs. Adopting existing ISA tools streamlines the implementation of our metadata checklist, providing user-friendly
interfaces for researchers to manage, document, and share animal phenotyping data efficiently.</p>]]></content><author><name>Sarah Oranna Fischer-Zielke</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[The current landscape of animal phenomics is characterised by a substantial lack of standardisation, hindering data reuse, reproducibility, and interoperability across studies, all of which are particularly important in light of the 3Rs principles for animal experiments (replace, reduce, refine). Within ELIXIR, the Domestic Animals Genome and Phenome Focus Group emerged to establish standardised practices that enhance the quality and interoperability of animal research data. In this context, the ISA model presents a robust, domain-agnostic framework well-established in the life sciences for describing experimental metadata. Notably, other scientific communities, such as the ELIXIR Plant and Metabolomics Communities (MIAPPE, PhenoMeNal), have successfully leveraged the ISA model to improve the consistency and usability of their metadata. Our project aims to develop a minimal information checklist tailored specifically for phenomics, facilitating the integration of diverse datasets, including recirculation systems in agriculture, and fostering collaborative research efforts. We will focus on various goals.Identifying essential aspects of animal phenotyping, informed by existing frameworks and community input. We aim to produce a concise and practical checklist that can be readily adopted by researchers, and promote a culture of standardisation.Mapping the checklist to the ISA model ensures alignment with established standards, promotes interoperability and facilitates data reuse while improving the overall quality of research outputs. Adopting existing ISA tools streamlines the implementation of our metadata checklist, providing user-friendly interfaces for researchers to manage, document, and share animal phenotyping data efficiently.]]></summary></entry><entry><title type="html">Evolving FAIR Image Analysis in Galaxy for Cross-domain and AI-ready Applications</title><link href="https://index.biohackrxiv.org//2026/03/31/tsxby.html" rel="alternate" type="text/html" title="Evolving FAIR Image Analysis in Galaxy for Cross-domain and AI-ready Applications" /><published>2026-03-31T00:00:00+00:00</published><updated>2026-03-31T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/03/31/tsxby</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/03/31/tsxby.html"><![CDATA[<p>The increasing adoption of image-based technologies across life sciences, environmental research, and related domains has increased the demand for
interoperable, reproducible, and FAIR-compliant image analysis infrastructures. At ELIXIR BioHackathon Europe 2025, Project 9, “Evolving FAIR Image
Analysis in Galaxy for Cross-domain and AI-ready Applications”, addressed these challenges by enhancing the Galaxy platform for bioimage analysis
with a focus on semantic interoperability, content-based reproducibility validation, and user-centered onboarding tutorials.To advance semantic
interoperability, we developed a curated vocabulary based on the EDAM Bioimaging ontology, which was applied to annotate tutorials on the Galaxy
Training Network, improving discoverability and aligning with evolving community standards. For reproducibility and AI-readiness, we integrated
the International Standard Content Code (ISCC) via the ISCC-SUM tool suite, enabling format-independent content-based validation, dataset
deduplication, and assessment of data similarity for robust model training. Finally, usability improvements included a comprehensive onboarding
tutorial for newcomers, enhanced integration with OMERO and BioImage Archive, and generally improved tool interoperability, including support for
GeoJSON-based spatial annotations. Collectively, these developments establish a scalable, cross-domain image analysis framework within Galaxy,
promoting FAIR-aligned practices while enabling reproducible and AI-ready workflows.</p>]]></content><author><name>Diana Chiang</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[The increasing adoption of image-based technologies across life sciences, environmental research, and related domains has increased the demand for interoperable, reproducible, and FAIR-compliant image analysis infrastructures. At ELIXIR BioHackathon Europe 2025, Project 9, “Evolving FAIR Image Analysis in Galaxy for Cross-domain and AI-ready Applications”, addressed these challenges by enhancing the Galaxy platform for bioimage analysis with a focus on semantic interoperability, content-based reproducibility validation, and user-centered onboarding tutorials.To advance semantic interoperability, we developed a curated vocabulary based on the EDAM Bioimaging ontology, which was applied to annotate tutorials on the Galaxy Training Network, improving discoverability and aligning with evolving community standards. For reproducibility and AI-readiness, we integrated the International Standard Content Code (ISCC) via the ISCC-SUM tool suite, enabling format-independent content-based validation, dataset deduplication, and assessment of data similarity for robust model training. Finally, usability improvements included a comprehensive onboarding tutorial for newcomers, enhanced integration with OMERO and BioImage Archive, and generally improved tool interoperability, including support for GeoJSON-based spatial annotations. Collectively, these developments establish a scalable, cross-domain image analysis framework within Galaxy, promoting FAIR-aligned practices while enabling reproducible and AI-ready workflows.]]></summary></entry><entry><title type="html">Tools to develop constraint-based models in R: adapting existing toolboxes</title><link href="https://index.biohackrxiv.org//2026/03/13/ey4c5.html" rel="alternate" type="text/html" title="Tools to develop constraint-based models in R: adapting existing toolboxes" /><published>2026-03-13T00:00:00+00:00</published><updated>2026-03-13T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/03/13/ey4c5</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/03/13/ey4c5.html"><![CDATA[<p>As part of the BioHackathon Europe 2025, we here report on the progress of the hacking team preparing tools to develop constraint-based
models in R for the Systems Biology community. This preliminary development relies on the adaptation of existing toolboxes. In this project,
we proposed the (re)development of an R based framework for developing and simulating constraint-based models. We proposed to expand
the Sybil library for model simulation with the functionalities for model reconstruction and analysis available in the widely used RAVEN
toolbox in Matlab. The outcome will facilitate constraint based modelling to experimental scientists, thereby contributing to bridge the
gap between data users and data generators. It will also be more FAIR by being usable with non-proprietary software, and align with
software best practices as collected by the ELIXIR Tools Platform. We will work towards increased reproducibility by also considering
implementation of FROG analysis in R. Moreover, as a tool developed by the ELIXIR Systems Biology Community for the wider community,
the long-term maintenance burden is spread across a wider membership.Two weeks before the BioHackathon, we discovered a new tool in R
allowing the simulation of models, called cobrar (https://github.com/Waschina/cobrar). Which calls for an assessment of its current
state and definition of new development areas.</p>]]></content><author><name>Jesubukade Ajakaye</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[As part of the BioHackathon Europe 2025, we here report on the progress of the hacking team preparing tools to develop constraint-based models in R for the Systems Biology community. This preliminary development relies on the adaptation of existing toolboxes. In this project, we proposed the (re)development of an R based framework for developing and simulating constraint-based models. We proposed to expand the Sybil library for model simulation with the functionalities for model reconstruction and analysis available in the widely used RAVEN toolbox in Matlab. The outcome will facilitate constraint based modelling to experimental scientists, thereby contributing to bridge the gap between data users and data generators. It will also be more FAIR by being usable with non-proprietary software, and align with software best practices as collected by the ELIXIR Tools Platform. We will work towards increased reproducibility by also considering implementation of FROG analysis in R. Moreover, as a tool developed by the ELIXIR Systems Biology Community for the wider community, the long-term maintenance burden is spread across a wider membership.Two weeks before the BioHackathon, we discovered a new tool in R allowing the simulation of models, called cobrar (https://github.com/Waschina/cobrar). Which calls for an assessment of its current state and definition of new development areas.]]></summary></entry><entry><title type="html">Bidirectional bridge: GitHub ⇄ bio.tools</title><link href="https://index.biohackrxiv.org//2026/02/24/8ktd6.html" rel="alternate" type="text/html" title="Bidirectional bridge: GitHub ⇄ bio.tools" /><published>2026-02-24T00:00:00+00:00</published><updated>2026-02-24T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/02/24/8ktd6</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/02/24/8ktd6.html"><![CDATA[<p>Research software metadata can be found across many code repositories and software registries. Here, we describe the tooling for a
bidirectional bridge between the software development platform GitHub and the ELIXIR bio.tools registry of life sciences software
tools and data resources. The developed bridge maps and improves metadata records across these two platforms, thereby benefiting
both and helping make research software more FAIR: findable, accessible, interoperable, and reusable. Specifically, the bridge
enables production of high-quality, rich bio.tools entries from the content already available in GitHub repositories, and uses
bio.tools records to suggest improvements to GitHub repositories through pull requests or issues. This includes adding missing
information and standardized descriptions for increased compliance with Software Management Plans. The bidirectional bridge makes
extensive use of existing APIs (GitHub, bio.tools, Europe PMC) and large language models (LLMs) to enrich metadata on both
platforms. By automating metadata extraction, improvement suggestion, and integration, the bridge reduces the manual overhead
required to FAIRify research software, lowering barriers for researchers to contribute or maintain well-annotated, reusable software.</p>]]></content><author><name>Mariia Steeghs-Turchina</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Research software metadata can be found across many code repositories and software registries. Here, we describe the tooling for a bidirectional bridge between the software development platform GitHub and the ELIXIR bio.tools registry of life sciences software tools and data resources. The developed bridge maps and improves metadata records across these two platforms, thereby benefiting both and helping make research software more FAIR: findable, accessible, interoperable, and reusable. Specifically, the bridge enables production of high-quality, rich bio.tools entries from the content already available in GitHub repositories, and uses bio.tools records to suggest improvements to GitHub repositories through pull requests or issues. This includes adding missing information and standardized descriptions for increased compliance with Software Management Plans. The bidirectional bridge makes extensive use of existing APIs (GitHub, bio.tools, Europe PMC) and large language models (LLMs) to enrich metadata on both platforms. By automating metadata extraction, improvement suggestion, and integration, the bridge reduces the manual overhead required to FAIRify research software, lowering barriers for researchers to contribute or maintain well-annotated, reusable software.]]></summary></entry><entry><title type="html">METRICS - Monitoring of Key Performance Indicators for ELIXIR Services</title><link href="https://index.biohackrxiv.org//2026/01/22/2jgk4.html" rel="alternate" type="text/html" title="METRICS - Monitoring of Key Performance Indicators for ELIXIR Services" /><published>2026-01-22T00:00:00+00:00</published><updated>2026-01-22T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2026/01/22/2jgk4</id><content type="html" xml:base="https://index.biohackrxiv.org//2026/01/22/2jgk4.html"><![CDATA[<p>Key Performance Indicators (KPIs) are increasingly requested by a diverse range of stakeholders across the research
ecosystem. Funders want to measure the impact of projects and related services they fund, or research organisations
want to track the service use for informed decision making. Service providers themselves are also interested in
monitoring their services to gather feedback and improve service quality. KPIs are a simple, but powerful tool for
these purposes.As part of the BioHackathon Europe 2025, we report on the activities of the METRICS project, which
addresses the need for consistent and transparent evaluation of services across ELIXIR and related initiatives
using KPIs. The project brings together experts from multiple ELIXIR Nodes and scientific domains to identify,
harmonise, and semantically model KPIs that reflect service quality, usage, sustainability, and impact. By exploring
existing evaluation frameworks, and processes, the team aims to design a flexible yet coherent foundation for KPI
monitoring of ELIXIR services. This report summarises the project’s motivation, current landscape analysis, and
initial steps toward developing an ontology-driven framework for KPI representation, fostering interoperability
and supporting evidence-based management of life science infrastructures.</p>]]></content><author><name>Nils-Christian Lübke</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Key Performance Indicators (KPIs) are increasingly requested by a diverse range of stakeholders across the research ecosystem. Funders want to measure the impact of projects and related services they fund, or research organisations want to track the service use for informed decision making. Service providers themselves are also interested in monitoring their services to gather feedback and improve service quality. KPIs are a simple, but powerful tool for these purposes.As part of the BioHackathon Europe 2025, we report on the activities of the METRICS project, which addresses the need for consistent and transparent evaluation of services across ELIXIR and related initiatives using KPIs. The project brings together experts from multiple ELIXIR Nodes and scientific domains to identify, harmonise, and semantically model KPIs that reflect service quality, usage, sustainability, and impact. By exploring existing evaluation frameworks, and processes, the team aims to design a flexible yet coherent foundation for KPI monitoring of ELIXIR services. This report summarises the project’s motivation, current landscape analysis, and initial steps toward developing an ontology-driven framework for KPI representation, fostering interoperability and supporting evidence-based management of life science infrastructures.]]></summary></entry><entry><title type="html">BioHackEU25 Report Project 16: MiCoReCa (Microbiome Community Resource Catalogue) - Towards Centralized Curation And Integration Of Microbiome Bioinformatics Resources</title><link href="https://index.biohackrxiv.org//2025/12/31/jfpsx.html" rel="alternate" type="text/html" title="BioHackEU25 Report Project 16: MiCoReCa (Microbiome Community Resource Catalogue) - Towards Centralized Curation And Integration Of Microbiome Bioinformatics Resources" /><published>2025-12-31T00:00:00+00:00</published><updated>2025-12-31T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2025/12/31/jfpsx</id><content type="html" xml:base="https://index.biohackrxiv.org//2025/12/31/jfpsx.html"><![CDATA[<p>The rapid growth of microbiome research has led to the development of numerous bioinformatics tools and databases, but information about them remains fragmented across disparate,
often outdated cataloging efforts, hindering resource discovery and utilization. To address this critical gap, the ELIXIR Microbiome Community proposes the development of MiCoReCa
(Microbiome Community Resource Catalogue), a comprehensive, dynamic, open-access catalogue of microbiome-related bioinformatics resources (tools, workflows, training, standards,
and databases). Leveraging our community’s expertise, this initiative will utilize standardized ontologies like EDAM and cross-reference established platforms like bio.tools and
WorkflowHub to create a centralized, findable inventory. A key feature is the community-driven process for identifying and curating missing ontological terms and metadata,
ensuring MiCoReCa’s accuracy and relevance in collaboration with partner platforms. Furthermore, the catalogue will integrate links to training materials from TeSS to support
appropriate tool usage, and connect with OpenEBench for benchmarking capabilities. This project will not only provide a vital resource for the microbiome field, enhancing
research efficiency and reproducibility, but will also establish a sustainable, adaptable infrastructure potentially applicable to other ELIXIR Communities. This effort
represents a significant contribution by the ELIXIR Microbiome Community to streamline microbiome bioinformatics.</p>]]></content><author><name>Vivek Ashokan</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[The rapid growth of microbiome research has led to the development of numerous bioinformatics tools and databases, but information about them remains fragmented across disparate, often outdated cataloging efforts, hindering resource discovery and utilization. To address this critical gap, the ELIXIR Microbiome Community proposes the development of MiCoReCa (Microbiome Community Resource Catalogue), a comprehensive, dynamic, open-access catalogue of microbiome-related bioinformatics resources (tools, workflows, training, standards, and databases). Leveraging our community’s expertise, this initiative will utilize standardized ontologies like EDAM and cross-reference established platforms like bio.tools and WorkflowHub to create a centralized, findable inventory. A key feature is the community-driven process for identifying and curating missing ontological terms and metadata, ensuring MiCoReCa’s accuracy and relevance in collaboration with partner platforms. Furthermore, the catalogue will integrate links to training materials from TeSS to support appropriate tool usage, and connect with OpenEBench for benchmarking capabilities. This project will not only provide a vital resource for the microbiome field, enhancing research efficiency and reproducibility, but will also establish a sustainable, adaptable infrastructure potentially applicable to other ELIXIR Communities. This effort represents a significant contribution by the ELIXIR Microbiome Community to streamline microbiome bioinformatics.]]></summary></entry><entry><title type="html">Decoding Complex Genotype-Phenotype Interactions by Discretizing the Genome</title><link href="https://index.biohackrxiv.org//2025/12/17/xhkc3.html" rel="alternate" type="text/html" title="Decoding Complex Genotype-Phenotype Interactions by Discretizing the Genome" /><published>2025-12-17T00:00:00+00:00</published><updated>2025-12-17T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2025/12/17/xhkc3</id><content type="html" xml:base="https://index.biohackrxiv.org//2025/12/17/xhkc3.html"><![CDATA[<p>Background: Despite the ease and affordability of genome sequencing in biomedical research, the genetic causes of many diseases
or their subtypes remain unknown due to diverse biological mechanisms that complicate genotype-phenotype relationships. Most
previous studies have focused on single variants or sets of variants presumed to be directly causal for the disease. However,
incomplete penetrance, in which some individuals carry disease-associated variants yet exhibit no phenotype, suggests that
these variants, the genomic background and other secondary factors combine to shape the susceptibility to the disease.</p>

<p>Results: Here, we introduce a new methodology for genotype-phenotype mapping based on genomic hashes, unique representations
of local genomic background. Each hash corresponds to a haplotype-resolved set of variants within one recombination-defined
genomic region (haploblock). We provide a practical guide for using genomic hashes to train machine learning models that link
genomic background and specific variant sets to phenotypic outcomes. We implemented this framework as a ready-to-use
bioinformatics pipeline capable of fast, scalable, hash-based genome comparison. The pipeline is available on GitHub:
https://github.com/collaborativebioinformatics/Haploblock_Clusters_ElixirBH25How it benefits the community: Genomic
hashes offer a computationally efficient framework for large-scale genotype-phenotype mapping. By discretizing the
genome into haploblocks, this approach will facilitate the search for causes of complex phenotypes across the entire
genome and the prediction of precision prevention points and treatments.</p>]]></content><author><name>Jędrzej Kubica</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Background: Despite the ease and affordability of genome sequencing in biomedical research, the genetic causes of many diseases or their subtypes remain unknown due to diverse biological mechanisms that complicate genotype-phenotype relationships. Most previous studies have focused on single variants or sets of variants presumed to be directly causal for the disease. However, incomplete penetrance, in which some individuals carry disease-associated variants yet exhibit no phenotype, suggests that these variants, the genomic background and other secondary factors combine to shape the susceptibility to the disease.]]></summary></entry><entry><title type="html">BioHackEU25 report: Towards a Robust Validation Service for Data and Metadata in ARC RO-Crates</title><link href="https://index.biohackrxiv.org//2025/12/16/zah28.html" rel="alternate" type="text/html" title="BioHackEU25 report: Towards a Robust Validation Service for Data and Metadata in ARC RO-Crates" /><published>2025-12-16T00:00:00+00:00</published><updated>2025-12-16T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2025/12/16/zah28</id><content type="html" xml:base="https://index.biohackrxiv.org//2025/12/16/zah28.html"><![CDATA[<p>Robust validation of both research data and its accompanying metadata is essential for ensuring adherence to FAIR principles.
Current approaches often handle these aspects separately, hindering a holistic quality assessment. Building upon previous
BioHackathon work establishing ARCs (Annotated Research Context) as RO-Crates (ARC RO-Crate), we aim to develop and demonstrate
an integrated validation strategy for FAIR digital objects. It distinguishes between validating the metadata descriptor and
the payload data files.For the metadata descriptor, validation will ensure structural and semantic compliance to the base
RO-Crate specification and the ARC-ISA family of RO-Crate profiles, using and extending the RO-Crate validator tool.For the
payload data files, validation targets the actual content, since data files often require domain-specific structural and
value constraints, which requires explicit schema definitions. For this, we will integrate Frictionless for checking data
content against community standards (e.g. MIAPPE, as demonstrated in the HORIZON project AGENT). Crucially, this project
will also explore mechanisms for specifying expected data structures’ requirements within the ARC RO-Crate itself. This
aims to provide a more self-contained description of data, investigating how such internal requirements can be linked to
data validation frameworks, complementing the crate’s metadata validation.The overall goal is to provide a powerful,
holistic validation mechanism for ARC RO-Crates, enhancing their reliability, trustworthiness, and FAIRness. A
MIAPPE-compliant plant phenomics dataset will serve as a use case. This integrated validation approach aims to streamline
quality control for researchers and will be packaged as a deployable microservice, offering broad applicability across
diverse research workflows.</p>]]></content><author><name>Eli Chadwick</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Robust validation of both research data and its accompanying metadata is essential for ensuring adherence to FAIR principles. Current approaches often handle these aspects separately, hindering a holistic quality assessment. Building upon previous BioHackathon work establishing ARCs (Annotated Research Context) as RO-Crates (ARC RO-Crate), we aim to develop and demonstrate an integrated validation strategy for FAIR digital objects. It distinguishes between validating the metadata descriptor and the payload data files.For the metadata descriptor, validation will ensure structural and semantic compliance to the base RO-Crate specification and the ARC-ISA family of RO-Crate profiles, using and extending the RO-Crate validator tool.For the payload data files, validation targets the actual content, since data files often require domain-specific structural and value constraints, which requires explicit schema definitions. For this, we will integrate Frictionless for checking data content against community standards (e.g. MIAPPE, as demonstrated in the HORIZON project AGENT). Crucially, this project will also explore mechanisms for specifying expected data structures’ requirements within the ARC RO-Crate itself. This aims to provide a more self-contained description of data, investigating how such internal requirements can be linked to data validation frameworks, complementing the crate’s metadata validation.The overall goal is to provide a powerful, holistic validation mechanism for ARC RO-Crates, enhancing their reliability, trustworthiness, and FAIRness. A MIAPPE-compliant plant phenomics dataset will serve as a use case. This integrated validation approach aims to streamline quality control for researchers and will be packaged as a deployable microservice, offering broad applicability across diverse research workflows.]]></summary></entry><entry><title type="html">Mining the potential of knowledge graphs for metadata on training</title><link href="https://index.biohackrxiv.org//2025/11/29/gv2ac.html" rel="alternate" type="text/html" title="Mining the potential of knowledge graphs for metadata on training" /><published>2025-11-29T00:00:00+00:00</published><updated>2025-11-29T00:00:00+00:00</updated><id>https://index.biohackrxiv.org//2025/11/29/gv2ac</id><content type="html" xml:base="https://index.biohackrxiv.org//2025/11/29/gv2ac.html"><![CDATA[<p>Training metadata in the life‑science community is increasingly standardized through Bioschemas, yet remains fragmented and under‑utilized. In
this work we harvested training records from ELIXR’s TeSS platform and the Galaxy Training Network, converting them into a unified knowledge
graph. A dedicated pipeline parses RDF/Turtle dumps, deduplicates entries, and builds rich indexes (keyword, provider, location, date, topic)
that power a Model Context Protocol (MCP) server. The MCP offers live and offline search tools—including keyword, provider, location, date,
topic, and SPARQL queries—enabling natural‑language access to training resources via LLM‑driven clients. User‑story driven evaluations
demonstrate the system’s ability to generate custom learning paths, assemble trainer profiles, and link training data to external repositories.
Findings highlight gaps in persistent identifiers (ORCID, ROR) and location granularity, informing recommendations for metadata providers. The
project showcases how knowledge‑graph‑backed metadata can enhance discoverability, interoperability, and AI‑assisted exploration of scientific
training materials.</p>]]></content><author><name>Dimitris Panouris</name></author><category term="BioHackEU25" /><summary type="html"><![CDATA[Training metadata in the life‑science community is increasingly standardized through Bioschemas, yet remains fragmented and under‑utilized. In this work we harvested training records from ELIXR’s TeSS platform and the Galaxy Training Network, converting them into a unified knowledge graph. A dedicated pipeline parses RDF/Turtle dumps, deduplicates entries, and builds rich indexes (keyword, provider, location, date, topic) that power a Model Context Protocol (MCP) server. The MCP offers live and offline search tools—including keyword, provider, location, date, topic, and SPARQL queries—enabling natural‑language access to training resources via LLM‑driven clients. User‑story driven evaluations demonstrate the system’s ability to generate custom learning paths, assemble trainer profiles, and link training data to external repositories. Findings highlight gaps in persistent identifiers (ORCID, ROR) and location granularity, informing recommendations for metadata providers. The project showcases how knowledge‑graph‑backed metadata can enhance discoverability, interoperability, and AI‑assisted exploration of scientific training materials.]]></summary></entry></feed>