<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator>
  <link href="https://index.biohackrxiv.org//feed.xml" rel="self" type="application/atom+xml"/>
  <link href="https://index.biohackrxiv.org/" rel="alternate" type="text/html"/>
  <updated>2026-09-04T15:28:29+00:00</updated>
  <id>https://index.biohackrxiv.org//archive.xml</id>
  <title type="html">BioHackrXiv Preprints</title>
  <subtitle>Preprints for BioHackathons</subtitle>
  <author>
    <name>BioHackrXiv</name>
    <uri>https://biohackrxiv.org/</uri>
  </author>

  
  <entry>
    <title type="html">AI-Assisted Variant Review Across Asia: Country-Level Expert Panels, Regional Collaboration, and Global Knowledge Sharing</title>
    <link href="https://index.biohackrxiv.org//2026/09/04/e5g6s.html" rel="alternate" type="text/html" title="AI-Assisted Variant Review Across Asia: Country-Level Expert Panels, Regional Collaboration, and Global Knowledge Sharing"/>
    <published>2026-09-04T00:00:00+00:00</published>
    <updated>2026-09-04T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/e5g6s_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/09/04/e5g6s.html">
      <![CDATA[ <p>Genome and exome sequencing have transformed rare disease diagnosis, yet converting large variant sets into evidence-backed interpretations remains labor-intensive
and fragmented. Across Asia, population genomic resources and specialist expertise are expanding, but prioritization, expert review, and reuse of reviewed knowledge
are often separated across institutions and countries. We argue that the most useful near-term role of artificial intelligence (AI) is not autonomous variant
classification, but reducing the friction between distributed evidence and distributed expert judgment. Building on collaborative platform development by participants
from institutions in Japan, Singapore, the Philippines, and Thailand, we propose a common workflow connecting variant prioritization and evidence organization,
structured expert review, and reviewed-knowledge sharing. AI can support multilingual phenotype structuring, population-aware candidate prioritization, literature
and evidence retrieval, and reuse of previous expert-reviewed records, while final evidence assessment remains expert-governed. Country-operated platforms can
preserve local governance and population context while exchanging standardized evidence and interpretations across Asia and contributing appropriate records to
global knowledge resources. Initial implementations for variant prioritization and an expert review workspace provide a practical foundation for this model. This
Perspective outlines where AI can add value, where human judgment must remain decisive, and how a regionally connected expert-panel network could be built.</p>

      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Toyofumi Fujiwara</name><uri>https://orcid.org/0000-0002-0170-9172</uri></author>
    
      <author><name>George Devasia</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Rutharra Ghayadthri Manisekaran</name><uri>https://orcid.org/0009-0008-6310-9974</uri></author>
    
      <author><name>Vasanthan Jayakumar</name><uri>https://orcid.org/0000-0002-6067-4184</uri></author>
    
      <author><name>Yosuke Kawai</name><uri>https://orcid.org/0000-0003-0666-1224</uri></author>
    
      <author><name>Shuichi Kawashima</name><uri>https://orcid.org/0000-0001-7883-3756</uri></author>
    
      <author><name>Yuko Kitano</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Francis A. Tablizo</name><uri>https://orcid.org/0009-0007-1392-3671</uri></author>
    
      <author><name>Shoichiro Takahashi</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Piyakrit Wongboonchai</name><uri>https://orcid.org/</uri></author>
    
    <category term="MHA26"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/09/04/e5g6s.html">
      <![CDATA[ Genome and exome sequencing have transformed rare disease diagnosis, yet converting large variant sets into evidence-backed interpretations remains labor-intensive and fragmented. Across Asia, population genomic resources and specialist expertise are expanding, but prioritization, expert review, and reuse of reviewed knowledge are often separated across institutions and countries. We argue that the most useful near-term role of artificial intelligence (AI) is not autonomous variant classification, but reducing the friction between distributed evidence and distributed expert judgment. Building on collaborative platform development by participants from institutions in Japan, Singapore, the Philippines, and Thailand, we propose a common workflow connecting variant prioritization and evidence organization, structured expert review, and reviewed-knowledge sharing. AI can support multilingual phenotype structuring, population-aware candidate prioritization, literature and evidence retrieval, and reuse of previous expert-reviewed records, while final evidence assessment remains expert-governed. Country-operated platforms can preserve local governance and population context while exchanging standardized evidence and interpretations across Asia and contributing appropriate records to global knowledge resources. Initial implementations for variant prioritization and an expert review workspace provide a practical foundation for this model. This Perspective outlines where AI can add value, where human judgment must remain decisive, and how a regionally connected expert-panel network could be built. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Data</title>
    <link href="https://index.biohackrxiv.org//2026/08/27/9mj78.html" rel="alternate" type="text/html" title="Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Data"/>
    <published>2026-08-27T00:00:00+00:00</published>
    <updated>2026-08-27T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/9mj78_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/27/9mj78.html">
      <![CDATA[ <p>As part of the BioHackathon Germany 2025, we report here about the progress of Project 8 -Enhancing e!DAL-PGP: A Modern Data Submission Platform
for Plant Science Research Dataduring the event.The increasing volume of data generated in plant research underscores the necessity for efficient
data management and sharing solutions. The de.NBI Service e!DAL-PGP (Arend et al., 2016, p. Arend2020) serves as a critical research data repository,
facilitating the storage, management, and dissemination of plant research data. However, the current implementation faces significant
challenges concerning the submission process and the provision of a submission tool for different operating systems, which complicate user
interactions and hinder data contribution. A primary issue with the existing e!DAL-PGP service is the cumbersome nature of maintaining and
deploying a submission tool across various OS environments. This requirement necessitates extensive effort to build, test, and provide the application
for each platform. Consequently, this fragmentation can lead to delays and inconsistencies in the submission process, ultimately hindering
researchers from effectively submitting their valuable data to the repository. To address these challenges, this project proposes the development
of a unified and user-friendly web submission tool that streamlines the data submission process to eliminate the complexities associated with
OS-specific requirements and to ensure that all users can submit their data seamlessly. This simplifies the submission process and enhances
usability by focussing onimproving the design and functionality. A well-structured and user-centric form is essential for facilitating accurate
and complete data submissions. The current interface lacks features that enhance user experience, such as lookup services, contextual help,
and clear instructions. By incorporating these elements, we aim to create a more efficient and engaging submission experience, encouraging
researchers to contribute their valuable data without unnecessary complexity. This initiative aligns closely with the goals of de.NBI, which
emphasizes the provision of high-quality bioinformatics services and the facilitation of FAIR Research Data Management (RDM). Enhancing the
e!DAL-PGP service will streamline the data submission process andpromote a culture of collaboration and data sharing within the plant research
community.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Manuel Feser</name><uri>https://orcid.org/0000-0001-6546-1818</uri></author>
    
      <author><name>Jonathan Bauer</name><uri>https://orcid.org/0000-0002-5624-2055</uri></author>
    
      <author><name>Sebastian Beier</name><uri>https://orcid.org/0000-0002-2177-8781</uri></author>
    
      <author><name>Dominik Brilhaus</name><uri>https://orcid.org/0000-0001-9021-3197</uri></author>
    
      <author><name>Dennis Psaroudakis</name><uri>https://orcid.org/0000-0002-7521-798X</uri></author>
    
      <author><name>Kevin Schneider</name><uri>https://orcid.org/0000-0002-2198-5262</uri></author>
    
      <author><name>Helena Schnitzer</name><uri>https://orcid.org/0000-0002-6382-9452</uri></author>
    
      <author><name>Heinrich Lukas Weil</name><uri>https://orcid.org/0000-0003-1945-6342</uri></author>
    
      <author><name>Daniel Arend</name><uri>https://orcid.org/0000-0002-2455-5938</uri></author>
    
    <category term="BH25DE"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/27/9mj78.html">
      <![CDATA[ As part of the BioHackathon Germany 2025, we report here about the progress of Project 8 -Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Dataduring the event.The increasing volume of data generated in plant research underscores the necessity for efficient data management and sharing solutions. The de.NBI Service e!DAL-PGP (Arend et al., 2016, p. Arend2020) serves as a critical research data repository, facilitating the storage, management, and dissemination of plant research data. However, the current implementation faces significant challenges concerning the submission process and the provision of a submission tool for different operating systems, which complicate user interactions and hinder data contribution. A primary issue with the existing e!DAL-PGP service is the cumbersome nature of maintaining and deploying a submission tool across various OS environments. This requirement necessitates extensive effort to build, test, and provide the application for each platform. Consequently, this fragmentation can lead to delays and inconsistencies in the submission process, ultimately hindering researchers from effectively submitting their valuable data to the repository. To address these challenges, this project proposes the development of a unified and user-friendly web submission tool that streamlines the data submission process to eliminate the complexities associated with OS-specific requirements and to ensure that all users can submit their data seamlessly. This simplifies the submission process and enhances usability by focussing onimproving the design and functionality. A well-structured and user-centric form is essential for facilitating accurate and complete data submissions. The current interface lacks features that enhance user experience, such as lookup services, contextual help, and clear instructions. By incorporating these elements, we aim to create a more efficient and engaging submission experience, encouraging researchers to contribute their valuable data without unnecessary complexity. This initiative aligns closely with the goals of de.NBI, which emphasizes the provision of high-quality bioinformatics services and the facilitation of FAIR Research Data Management (RDM). Enhancing the e!DAL-PGP service will streamline the data submission process andpromote a culture of collaboration and data sharing within the plant research community. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">BioHackSWAT4HCLS25 report: Towards an interactive mapping experience for data owners</title>
    <link href="https://index.biohackrxiv.org//2026/08/17/mhvbs.html" rel="alternate" type="text/html" title="BioHackSWAT4HCLS25 report: Towards an interactive mapping experience for data owners"/>
    <published>2026-08-17T00:00:00+00:00</published>
    <updated>2026-08-17T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/mhvbs_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/17/mhvbs.html">
      <![CDATA[ <p>At the Barcelona SWAT4HCLS 2025 Hackathon, a hacking group familiarized with and worked on improvements for RDFCraft. A tool for a data
onboarding tool that helps with mapping tabular or JSON formatted data to a reference schema ontology.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Karolis Cremers</name><uri>https://orcid.org/0000-0002-1756-3905</uri></author>
    
      <author><name>Daphne Wijnbergen</name><uri>https://orcid.org/0000-0002-7449-6657</uri></author>
    
      <author><name>Julia Koblitz</name><uri>https://orcid.org/0000-0002-7260-2129</uri></author>
    
      <author><name>Javier Millan Acosta</name><uri>https://orcid.org/0000-0002-4166-7093</uri></author>
    
      <author><name>Rick Overkleeft</name><uri>https://orcid.org/0009-0004-3529-1159</uri></author>
    
      <author><name>David Schimmel</name><uri>https://orcid.org/0000-0002-1719-8928</uri></author>
    
      <author><name>Katja Hoffmann</name><uri>https://orcid.org/0000-0003-4765-0767</uri></author>
    
      <author><name>Brieuc Quemeneur</name><uri>https://orcid.org/0009-0006-7574-2254</uri></author>
    
      <author><name>Ensar Emir Erol</name><uri>https://orcid.org/0000-0002-5739-8860</uri></author>
    
    <category term="SWAT4HCLS25"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/17/mhvbs.html">
      <![CDATA[ At the Barcelona SWAT4HCLS 2025 Hackathon, a hacking group familiarized with and worked on improvements for RDFCraft. A tool for a data onboarding tool that helps with mapping tabular or JSON formatted data to a reference schema ontology. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">INTOXICOM Workshop Report: Making toxicology tools more accessible and interoperable</title>
    <link href="https://index.biohackrxiv.org//2026/08/14/u3fnh.html" rel="alternate" type="text/html" title="INTOXICOM Workshop Report: Making toxicology tools more accessible and interoperable"/>
    <published>2026-08-14T00:00:00+00:00</published>
    <updated>2026-08-14T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/u3fnh_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/14/u3fnh.html">
      <![CDATA[ <p>As part of the INTOXICOM Implementation Study for the ELIXIR Toxicology Community a series of workshops is organized (Martens et al., 2024). Here,
we here report on the 3rd workshop, titled “Making toxicology tools more accessible and interoperable” which was held from 26 to 27 March 2025 at
the SciLifeLab at Uppsala University in Sweden. The workshop welcomed 29 participants from Sweden, The Netherlands, Cyprus, Switzerland, Italy,
Greece, France, and Norway. This 3rd INTOXICOM workshop covered various aspects of making computational tools available and how they are used.
Several projects, such as NanoSolveIT (Afantitis et al., 2020), ONTOX (Vinken et al., 2021), and VHP4Safety (Kienhuis et al., 2024), already make
computational toxicology services available, but the field of integrating computional toxicology dates back much longer, such as Bioclipse developed
at Uppsala University (Willighagen et al., 2011), making it a perfect location to have held this workshop.</p>

      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Meike Bünger</name><uri>https://orcid.org/0009-0002-7664-0058</uri></author>
    
      <author><name>Ola Spjuth</name><uri>https://orcid.org/0000-0002-8083-2864</uri></author>
    
      <author><name>Nikita Churikov</name><uri>https://orcid.org/0009-0001-5548-0561</uri></author>
    
      <author><name>Jonne Rietdijk</name><uri>https://orcid.org/0000-0003-3799-6684</uri></author>
    
      <author><name>Jente Houweling</name><uri>https://orcid.org/0009-0005-3680-0645</uri></author>
    
      <author><name>Wolmar Nyberg Åkerström</name><uri>https://orcid.org/0000-0002-3890-6620</uri></author>
    
      <author><name>Ogunleye Adeolu</name><uri>https://orcid.org/0000-0001-6763-4807</uri></author>
    
      <author><name>Penny Nymark</name><uri>https://orcid.org/0000-0002-3435-7775</uri></author>
    
      <author><name>Petru Niga</name><uri>https://orcid.org/0000-0003-0195-3850</uri></author>
    
      <author><name>Cleo Tebby</name><uri>https://orcid.org/0000-0003-3470-157X</uri></author>
    
      <author><name>Egon Willighagen</name><uri>https://orcid.org/0000-0001-7542-0286</uri></author>
    
    <category term="INTOXICOM"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/14/u3fnh.html">
      <![CDATA[ As part of the INTOXICOM Implementation Study for the ELIXIR Toxicology Community a series of workshops is organized (Martens et al., 2024). Here, we here report on the 3rd workshop, titled “Making toxicology tools more accessible and interoperable” which was held from 26 to 27 March 2025 at the SciLifeLab at Uppsala University in Sweden. The workshop welcomed 29 participants from Sweden, The Netherlands, Cyprus, Switzerland, Italy, Greece, France, and Norway. This 3rd INTOXICOM workshop covered various aspects of making computational tools available and how they are used. Several projects, such as NanoSolveIT (Afantitis et al., 2020), ONTOX (Vinken et al., 2021), and VHP4Safety (Kienhuis et al., 2024), already make computational toxicology services available, but the field of integrating computional toxicology dates back much longer, such as Bioclipse developed at Uppsala University (Willighagen et al., 2011), making it a perfect location to have held this workshop. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">4th BioHackathon Germany report: Exploring Gamification Strategies to Enhance Bioinformatics Training</title>
    <link href="https://index.biohackrxiv.org//2026/08/10/dfwm9.html" rel="alternate" type="text/html" title="4th BioHackathon Germany report: Exploring Gamification Strategies to Enhance Bioinformatics Training"/>
    <published>2026-08-10T00:00:00+00:00</published>
    <updated>2026-08-10T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/dfwm9_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/10/dfwm9.html">
      <![CDATA[ <p>State of the art life science training features steep learning curves due to dense technical specifications and complex data formats,
often causing cognitive overload and low learner retention. While gamification can enhance engagement, implementing it without
trivializing scientific content remains challenging. As part of the Biohackathon Germany 2025, we explored strategies to adapt
gamification for bioinformatics education. We curated a resource matrix evaluating 20 digital tools based on cost, implementation
effort, and pedagogical impact. To guide instructors, we formulated the “Ten Simple Rules for Gamification in Bioinformatics and
Life Sciences Education,” emphasizing a shift from superficial point systems to deep, competency-driven mechanics rooted in authentic
data and high-stakes narratives. We validated this framework through two pilot implementations: translating an introductory R
programming course into interactive console tutorials using swirl accelerated by Large Language Models (LLMs), and deploying
browser-based Research Data Management (RDM) quizzes via Wordwall to reinforce FAIR principles. Our findings reveal that while
specialized tools fit specific niches easily, broader open-source frameworks offer greater flexibility, with implementation workloads
significantly mitigated by generative AI workflows. Ultimately, gamification serves as a powerful pedagogical asset when balanced
correctly, transforming abstract computational workflows into engaging, collaborative simulations that bridge virtual training and
professional scientific competency.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Dominik Lux</name><uri>https://orcid.org/0000-0002-7490-8260</uri></author>
    
      <author><name>Daniel Wibberg</name><uri>https://orcid.org/0000-0002-1331-4311</uri></author>
    
      <author><name>Martin Eisenacher</name><uri>https://orcid.org/0000-0003-2687-7444</uri></author>
    
      <author><name>Karin Schork</name><uri>https://orcid.org/0000-0003-3756-4347</uri></author>
    
    <category term="BH25DE"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/10/dfwm9.html">
      <![CDATA[ State of the art life science training features steep learning curves due to dense technical specifications and complex data formats, often causing cognitive overload and low learner retention. While gamification can enhance engagement, implementing it without trivializing scientific content remains challenging. As part of the Biohackathon Germany 2025, we explored strategies to adapt gamification for bioinformatics education. We curated a resource matrix evaluating 20 digital tools based on cost, implementation effort, and pedagogical impact. To guide instructors, we formulated the “Ten Simple Rules for Gamification in Bioinformatics and Life Sciences Education,” emphasizing a shift from superficial point systems to deep, competency-driven mechanics rooted in authentic data and high-stakes narratives. We validated this framework through two pilot implementations: translating an introductory R programming course into interactive console tutorials using swirl accelerated by Large Language Models (LLMs), and deploying browser-based Research Data Management (RDM) quizzes via Wordwall to reinforce FAIR principles. Our findings reveal that while specialized tools fit specific niches easily, broader open-source frameworks offer greater flexibility, with implementation workloads significantly mitigated by generative AI workflows. Ultimately, gamification serves as a powerful pedagogical asset when balanced correctly, transforming abstract computational workflows into engaging, collaborative simulations that bridge virtual training and professional scientific competency. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Variant representation in RDF</title>
    <link href="https://index.biohackrxiv.org//2026/08/09/jazsb.html" rel="alternate" type="text/html" title="Variant representation in RDF"/>
    <published>2026-08-09T00:00:00+00:00</published>
    <updated>2026-08-09T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/jazsb_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/09/jazsb.html">
      <![CDATA[ <p>During the International SWAT4HCLS conference held on 24-27th February 2025 in Barcelona (Spain), we detected an emerging number of novel RDF models to
represent variant information in genomic datasets potentially hindering data reuse. We tackled the question how semantic representations can enhance the
interoperability of variant data for clinical applications. Here we report our initial results on genomic variant schema alignment.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Núria Queralt-Rosinach</name><uri>https://orcid.org/0000-0003-0169-8159</uri></author>
    
      <author><name>Alexander Jonathan Kellmann</name><uri>https://orcid.org/0000-0001-6108-5552</uri></author>
    
      <author><name>Alexandrina Bodrug-Schepers</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Alban Gaignard</name><uri>https://orcid.org/0000-0002-3597-8557</uri></author>
    
      <author><name>Elias Crum</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Pierre Larmande</name><uri>https://orcid.org/</uri></author>
    
      <author><name>Andra Waagmeester</name><uri>https://orcid.org/0000-0001-9773-4008</uri></author>
    
      <author><name>Jerven Bolleman</name><uri>https://orcid.org/0000-0002-7449-1266</uri></author>
    
    <category term="SWAT4HCLS25"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/09/jazsb.html">
      <![CDATA[ During the International SWAT4HCLS conference held on 24-27th February 2025 in Barcelona (Spain), we detected an emerging number of novel RDF models to represent variant information in genomic datasets potentially hindering data reuse. We tackled the question how semantic representations can enhance the interoperability of variant data for clinical applications. Here we report our initial results on genomic variant schema alignment. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Variant annotation in RDF for clinical trials matching</title>
    <link href="https://index.biohackrxiv.org//2026/08/09/2hbqf.html" rel="alternate" type="text/html" title="Variant annotation in RDF for clinical trials matching"/>
    <published>2026-08-09T00:00:00+00:00</published>
    <updated>2026-08-09T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/2hbqf_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/08/09/2hbqf.html">
      <![CDATA[ <p>Precision oncology depends on semantic, interoperable representations of genomic variants (GV) - particularly structural variants (SVs) - to match patients
with clinical trials. In this exploratory project, we investigated the use of RDF and the GA4GH VRS Schema to standardize variant annotations and integrate
them with clinical trial data. Our work, developed in collaboration with the Pangenome Graphs and Platform for Precision Medicine groups, prototypes an
RDF-based data harmonization that paves the way for improved semantic interoperability in precision medicine, especially for cancer research and AI-driven
discovery.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Núria Queralt-Rosinach</name><uri>https://orcid.org/0000-0003-0169-8159</uri></author>
    
      <author><name>Toshiaki Katayama</name><uri>https://orcid.org/0000-0003-2391-0384</uri></author>
    
      <author><name>Jose Emilio Labra-Gayo</name><uri>https://orcid.org/0000-0001-8907-5348</uri></author>
    
      <author><name>Claude Nanjo</name><uri>https://orcid.org/0009-0002-1208-8858</uri></author>
    
    <category term="BH25JP"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/08/09/2hbqf.html">
      <![CDATA[ Precision oncology depends on semantic, interoperable representations of genomic variants (GV) - particularly structural variants (SVs) - to match patients with clinical trials. In this exploratory project, we investigated the use of RDF and the GA4GH VRS Schema to standardize variant annotations and integrate them with clinical trial data. Our work, developed in collaboration with the Pangenome Graphs and Platform for Precision Medicine groups, prototypes an RDF-based data harmonization that paves the way for improved semantic interoperability in precision medicine, especially for cancer research and AI-driven discovery. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Schema-Driven Generation of Synthetic HL7 FHIR RDF Data from Shape Expressions (ShEx)</title>
    <link href="https://index.biohackrxiv.org//2026/07/29/3gak2.html" rel="alternate" type="text/html" title="Schema-Driven Generation of Synthetic HL7 FHIR RDF Data from Shape Expressions (ShEx)"/>
    <published>2026-07-29T00:00:00+00:00</published>
    <updated>2026-07-29T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/3gak2_v2</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/07/29/3gak2.html">
      <![CDATA[ <p>We describe how synthetic HL7 FHIR data in RDF was produced directly from Shape Expressions (ShEx), using the authoritative FHIR R4 ShEx schema as the
sole source of domain structure. Rather than encoding clinical knowledge in a domain-specific simulator, we drive generation from the published shapes:
a schema-driven generator (rudof generate) consumes them, and a small configuration file controls scale, cardinalities, and value generation. This note
reports the method - schema selection, a minimal schema preparation step, the generator configuration, and the invocation - so that the process is
reproducible.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Diego Martín Fernández</name><uri>https://orcid.org/0009-0003-6640-9474</uri></author>
    
      <author><name>Álvaro García Fernández</name><uri>https://orcid.org/0009-0008-9390-6210</uri></author>
    
      <author><name>Samuel Bustamante Larriet</name><uri>https://orcid.org/0009-0005-8631-2682</uri></author>
    
      <author><name>Eric Prud'hommeaux</name><uri>https://orcid.org/0000-0003-1775-9921</uri></author>
    
      <author><name>Jose Emilio Labra-Gayo</name><uri>https://orcid.org/0000-0001-8907-5348</uri></author>
    
    <category term="GOBLINHack26"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/07/29/3gak2.html">
      <![CDATA[ We describe how synthetic HL7 FHIR data in RDF was produced directly from Shape Expressions (ShEx), using the authoritative FHIR R4 ShEx schema as the sole source of domain structure. Rather than encoding clinical knowledge in a domain-specific simulator, we drive generation from the published shapes: a schema-driven generator (rudof generate) consumes them, and a small configuration file controls scale, cardinalities, and value generation. This note reports the method - schema selection, a minimal schema preparation step, the generator configuration, and the invocation - so that the process is reproducible. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Measure before you rewrite: ablation-driven redesign of LLM-facing RDF schema documentation in TogoMCP</title>
    <link href="https://index.biohackrxiv.org//2026/07/26/6v5ra.html" rel="alternate" type="text/html" title="Measure before you rewrite: ablation-driven redesign of LLM-facing RDF schema documentation in TogoMCP"/>
    <published>2026-07-26T00:00:00+00:00</published>
    <updated>2026-07-26T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/6v5ra_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/07/26/6v5ra.html">
      <![CDATA[ <p>MIE files are per-database YAML documents that TogoMCP supplies to a large language model at  query time so it can compose SPARQL against the
DBCLS RDF Portal. Ours had grown to eleven  sections, semi-automatically generated for each of 36 databases and reviewed by hand. Each  section
had been introduced in response to a systematic query failure observed in use — sound  practice, but it left open whether any section still
earned its tokens once the other ten were  present. We measured that, in eighteen ablation conditions across four families:  is a section
necessary, is a functional group necessary, is the whole document worth anything,  and is any one group sufficient alone. No single section
and no single group is necessary.  Removing the entire document costs 0.9 points out of 20, and the query-construction group alone  recovers
99% of that — the whole is worth roughly 2.7 times the sum of its parts, the signature  of heavy redundancy. We rebuilt the format around
that evidence, making the verified executable  example the atomic unit: 36 files, 303 examples, each 29–65% smaller than the file it replaces.
A pre-registered equivalence run over 100 benchmark questions finds v3 statistically  indistinguishable from v2 in answer quality (+0.29/20,
95% CI [-0.09, +0.68]) while using 15%  fewer input tokens, costing 15% less and running 6% faster, with the factoid-question score up  a
full point. We also report eight measurement traps that faked or destroyed signal, and one  budgeting error worth more than the results:
we spent the most on the least informative  experiment, and say what we would do instead.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Akira R Kinjo</name><uri>https://orcid.org/0000-0002-4006-8208</uri></author>
    
      <author><name>Yasunori Yamamoto</name><uri>https://orcid.org/0000-0002-6943-6887</uri></author>
    
    <category term="BH25JP"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/07/26/6v5ra.html">
      <![CDATA[ MIE files are per-database YAML documents that TogoMCP supplies to a large language model at query time so it can compose SPARQL against the DBCLS RDF Portal. Ours had grown to eleven sections, semi-automatically generated for each of 36 databases and reviewed by hand. Each section had been introduced in response to a systematic query failure observed in use — sound practice, but it left open whether any section still earned its tokens once the other ten were present. We measured that, in eighteen ablation conditions across four families: is a section necessary, is a functional group necessary, is the whole document worth anything, and is any one group sufficient alone. No single section and no single group is necessary. Removing the entire document costs 0.9 points out of 20, and the query-construction group alone recovers 99% of that — the whole is worth roughly 2.7 times the sum of its parts, the signature of heavy redundancy. We rebuilt the format around that evidence, making the verified executable example the atomic unit: 36 files, 303 examples, each 29–65% smaller than the file it replaces. A pre-registered equivalence run over 100 benchmark questions finds v3 statistically indistinguishable from v2 in answer quality (+0.29/20, 95% CI [-0.09, +0.68]) while using 15% fewer input tokens, costing 15% less and running 6% faster, with the factoid-question score up a full point. We also report eight measurement traps that faked or destroyed signal, and one budgeting error worth more than the results: we spent the most on the least informative experiment, and say what we would do instead. ]]>
    </summary></entry>
  
  <entry>
    <title type="html">Maintaining and refining the Tidyomics ecosystem: enhancing core packages and interoperability for EuroBioc2026</title>
    <link href="https://index.biohackrxiv.org//2026/07/08/cd9s6.html" rel="alternate" type="text/html" title="Maintaining and refining the Tidyomics ecosystem: enhancing core packages and interoperability for EuroBioc2026"/>
    <published>2026-07-08T00:00:00+00:00</published>
    <updated>2026-07-08T00:00:00+00:00</updated>
    <id>https://doi.org/10.37044/osf.io/cd9s6_v1</id>
    <content type="html" xml:base="https://index.biohackrxiv.org//2026/07/08/cd9s6.html">
      <![CDATA[ <p>The Tidyomics ecosystem facilitates the manipulation of computational omics data structures by bringing the intuitive and
consistent syntax of the tidy paradigm to R. During the EuroBioc2026 Tidyomics Hackathon, five bioinformatics researchers
collaborated to strengthen this ecosystem across four areas. First, we introduce tidyAnnData, a new package that expands
interoperability between the tidyverse and AnnData objects. Second, we updated and harmonized the accessibility of
information across the core packages that form the current Tidyomics backbone. Third, we improved the stability of core
packages by resolving critical bugs through targeted pull requests and implementing functional enhancements to the DFplyr
and tidybulk packages. Fourth, we enhanced the documentation by producing a comprehensive and stable vignette for
tidySingleCellExperiment covering typical single-cell analysis workflows. Together, these contributions lower the barrier
to entry for new users, promote reproducibility, and support the continued transition from disparate scripts toward robust,
unified omics workflows driven by community development.</p>


      <h4>References</h4>
      <ul>
      </ul>
      ]]>
    </content>
    
    
      <author><name>Carissa Chen</name><uri>https://orcid.org/0000-0002-9225-7086</uri></author>
    
      <author><name>Marco Geigges</name><uri>https://orcid.org/0000-0001-9071-5162</uri></author>
    
      <author><name>Jasper Spitzer</name><uri>https://orcid.org/0000-0001-9696-2092</uri></author>
    
      <author><name>Stevie Pederson</name><uri>https://orcid.org/0000-0001-8197-3303</uri></author>
    
      <author><name>Juan Henao</name><uri>https://orcid.org/0000-0003-0783-1432</uri></author>
    
    <category term="EuroBioc2026"/>
    
    <summary type="html" xml:base="https://index.biohackrxiv.org//2026/07/08/cd9s6.html">
      <![CDATA[ The Tidyomics ecosystem facilitates the manipulation of computational omics data structures by bringing the intuitive and consistent syntax of the tidy paradigm to R. During the EuroBioc2026 Tidyomics Hackathon, five bioinformatics researchers collaborated to strengthen this ecosystem across four areas. First, we introduce tidyAnnData, a new package that expands interoperability between the tidyverse and AnnData objects. Second, we updated and harmonized the accessibility of information across the core packages that form the current Tidyomics backbone. Third, we improved the stability of core packages by resolving critical bugs through targeted pull requests and implementing functional enhancements to the DFplyr and tidybulk packages. Fourth, we enhanced the documentation by producing a comprehensive and stable vignette for tidySingleCellExperiment covering typical single-cell analysis workflows. Together, these contributions lower the barrier to entry for new users, promote reproducibility, and support the continued transition from disparate scripts toward robust, unified omics workflows driven by community development. ]]>
    </summary></entry>
  
</feed>
