RSAA26 Conference Schedule & Abstracts
Day 1 – Tuesday 25 August 2026
Community, Collaboration, Skills, Training & Careers
Time: 12:30 – 17:15 AEST (UTC+10)
| 12:30–13:15 | Welcome & Conference Introduction (45 mins) | ||
| 13:15–13:45 | Day 1 Keynote Presenter: Liana Jacinta (30 mins) | ||
| Time | Session 1: Diversity & Policy Panel | Session 2: Equity, Access & Knowledge | Session 3: Community & Collaboration |
|---|---|---|---|
| 13:45–15:25 (100 mins) | • 13:45–14:45 (60 mins): Making space: Celebrating people and diversity in Research Software Presenter: Jana Makar (Women in HPC) • 14:45–15:25 (40 mins): The SSI Fellowship Without Borders: What Would Work, and What Would Not Presenters: Oscar Seip, Saranjeet Kaur, Aman Goel | Interactive Workshop: Meta-valuation Across Borders: A Cross-Community Experiment in Valuing Diverse Research Contributions Presenter: Dr Cooper Smout (Runs 13:45–15:15, 90 mins) | Birds of a Feather (BoF): BoF: Metadata Across Disciplines – A Collaborative Journey Presenters: Paul Wang, Emily Fitzgerald, Scout Bell, Rowland Mosbergen (runs 13:45–15:15, 90 mins) |
| 15:25–15:35 | Transition Break (10 mins) | ||
| Time | Session 4: Community & Collaboration | Session 5: Community & Collaboration | Session 6: Sustainability & Impact |
| 15:35–17:05 (90 mins) | Interactive BoF: Bridging the Gap: BoF for Standardising User and Dataset Mapping Across Research Software Ecosystems Presenters: Rowland Mosbergen, Anurag Katariya, Melroy Almeida, Sarah Thomas (Australian Access Federation) (Runs 15:35–16:35, 60 mins) | Birds of a Feather: Birds of a Feather: RSE-AUNZ Update, Community Survey & Ask Me Anything Presenters: Rowland Mosbergen, Matthew Laurenson (90 mins) | Mini Workshop: Computational Research Sustainability Presenter: Pao Corrales (Runs 15:35–16:35, 60 mins) |
| 17:05–17:10 | Break (5 mins) | ||
| 17:10–17:15 | Day 1 Closing & Announcements (5 mins) | ||
Day 2 – Wednesday 26 August 2026
Skills, Training, Careers, Reproducibility, FAIR, Open Research, AI & Emerging Tools
Time: 12:30 – 17:00 AEST (UTC+10)
| 12:30–12:45 | Day 2 Welcome & Announcements (15 mins) | |
| 12:45–13:15 | Day 2 Keynote From method to research software, and beyond: 18 years of building mixOmics Presenter: Prof. Kim-Anh Lê Cao (30 mins) | |
| Time | Session 7: Skills, Training & Careers | Session 8: Reproducibility, FAIR & Open Research |
|---|---|---|
| 13:15–14:30 (75 mins) | • 13:15–13:30 (15 mins): The RCP Discovery Internship Program: Building a Global Innovation Lab for Research Software Presenter: Rowland Mosbergen • 13:30–13:45 (15 mins): The Future (and Present) of Research Software Engineering with Cross-functional Teams Presenter: Gregory B. Poole • 13:45–14:00 (15 mins): Fortran is cool don't let your friends convince you otherwise Presenter: Jorge Luis Galvez Vallejo • 14:00–14:15 (15 mins): From Ad Hoc Promotions to Visible Career Pathways for Research Software Engineers Presenters: Rowland Mosbergen, Julie Iskander • 14:15–14:30 (15 mins): Building Capacity from the Inside: A Project Update from the Research Software Engineering Capacity Enhancement Project (RSE-CEP) Presenter: James Smithies | • 13:15–13:30 (15 mins): An Integrative Bioinformatics Framework for Identifying Core Regulatory Pathways in Melanoma Metastasis Presenters: Tifany Nabilah, Linda Erlina • 13:30–13:45 (15 mins): A Reproducible Bioinformatics Workflow for Biomarker Discovery in Therapy-Resistant Breast Cancer Using Public Transcriptomic Datasets Presenter: Sepiah Dwi Cahyani • 13:45–14:00 (15 mins): Computational Transcriptomics Analysis for Biomarker Discovery in Triple Negative Breast Cancer Presenter: May Eva Francisca • 14:00–14:15 (15 mins): Scaling Reproducible Bioinformatics Workflows with R Markdown & Slurm Presenter: Felicia Ong Sing Yi • 14:15–14:30 (15 mins): From R to Python: Porting a Research Software Package Across Language Ecosystems Presenter: Jyoti Bhogal |
| 14:30–14:45 | Session 7 Question Time & Discussion (15 mins) | Session 8 Question Time & Discussion (15 mins) |
| 14:45–15:15 | Session 7 Breakout & Chat Rooms (30 mins) | Session 8 Breakout & Chat Rooms (30 mins) |
| 15:15–15:30 | Coffee Break / Screen-Off Eye Break (15 mins) | |
| Time | Session 9: AI & Emerging Tools | Session 10: Reproducibility, FAIR & Open Research |
| 15:30–16:15 (45 mins) | Demos & Showcases (10 mins): • 15:30–15:40 (10 mins): Namo: Working with Named Dimensional Data Presenter: thoran • 15:40–15:50 (10 mins): BRI DataLab: an AI-Assisted Research Platform for Infrastructure Finance Research in International Relations Presenter: Sanoop Sajan Koshy Lightning Talks (5 mins): • 15:50–15:55 (5 mins): Research Software Engineering in the Age of Generative AI Presenters: Richard Littauer, Michelle Barker, Sandra Gesing • 15:55–16:00 (5 mins): Breaking Bias in Requirements Engineering: Artificial Intelligence for Inclusivity & Sustainability Presenter: Dr Hafsa Shareef Dar • 16:00–16:05 (5 mins): The HelpfulBatBot: a RAG-based assistant for democratising Geodynamic modelling Presenters: Juan Carlos Graciosa, Louis Moresi • 16:05–16:10 (5 mins): An Efficient High Dimensional Gene Expression Based Multi-Class Cancer Prediction Model Using Ensemble Learning Presenter: P. Krithy Sreshta • 16:10–16:15 (5 mins): Intelligent Automation in Healthcare Innovation Presenter: Anushka Bhargava | Demos & Showcases (10 mins): • 15:30–15:40 (10 mins): HPC Resource Dashboard: Helping Researchers Do More With Less on Shared HPC Infrastructure Presenter: Samuel Green Lightning Talks (5 mins): • 15:40–15:45 (5 mins): Reproducible command-line workflows with asciinema-scripted Presenter: Robert Moss • 15:45–15:50 (5 mins): Breaking Barriers in Research Data: Implementing FAIR Workflows and Persistent Identifiers Across Distributed Systems Presenter: Aline de Souza Andrade • 15:50–15:55 (5 mins): Ecosystem Indicators Workflows: Building a FAIR, Modular Toolbox for Reproducible Ecological Assessment in Australia Presenters: Abhimanyu Raj Singh, Jose Ferrer Paris • 15:55–16:00 (5 mins): Got a Data Challenge? There's a Global Community for That: An Introduction to the Research Data Alliance (RDA) Presenter: Trish Radotic • 16:00–16:05 (5 mins): A 3-month challenge to revive, share, and sustain a PhD-era bioinformatics tool Presenter: Saliha Elif Yildizhan • 16:05–16:10 (5 mins): Life Registries and the Unborn Child: A Matter of Data Access and Protection? Presenter: Dennis Batangan • 16:10–16:15 (5 mins): Building and sustaining research software engineering communities: The role of computational skills trainers Presenter: Elijah Atta Manu |
| 16:15–16:30 | Session 9 Question Time & Discussion (15 mins) | Session 10 Question Time & Discussion (15 mins) |
| 16:30–16:55 | Session 9 Breakout & Chat Rooms (25 mins) | Session 10 Breakout & Chat Rooms (25 mins) |
| 16:55–17:00 | Day 2 Closing & Announcements (5 mins) | |
Day 3 – Thursday 27 August 2026
Infrastructure, Platforms, Skills, Careers, Data Platforms, Data Spaces, Equity & Access
Time: 12:30 – 17:00 AEST (UTC+10)
| 12:30–12:40 | Day 3 Welcome & Daily Overview (10 mins) | |
| Time | Session 11: Infrastructure & Platforms | Session 12: Equity, Access & Knowledge |
|---|---|---|
| 12:40–13:55 (75 mins) | • 12:40–12:55 (15 mins): Launching Discworld: A bespoke, user-considerate, CryoEM processing cluster Presenters: Nathan Glades, Keiran Rowell • 12:55–13:10 (15 mins): CPUX: State-Driven Orchestration for Distributed Research Simulations Presenter: Pronab Pal • 13:10–13:25 (15 mins): In the Beginning Was Documentation, and the Documentation Was Parsed, and the Documentation Begat the Code, and It Was Good Presenter: Elio Campitelli • 13:25–13:40 (15 mins): Australian BioCommons: infrastructure access and execution for complex research workflows Presenters: Johan Gustafsson, Steven Manos | • 12:40–12:55 (15 mins): Meta-valuation: A Self-Referential Mechanism for Valuing Diverse Research Contributions Presenter: Dr Cooper Smout • 12:55–13:10 (15 mins): From Archive to Analysis: Community-Led Open Software for HASS, Indigenous and GLAM Digital Collections Presenter: Natalia Orel • 13:10–13:25 (15 mins): Dual-Purpose Research Software Presenter: Attila Egri-Nagy • 13:25–13:40 (15 mins): Open Source as Public Infrastructure: Early Evidence from Africa Presenters: Tosan Okome, Chidera Frankie • 13:40–13:55 (15 mins): Software Without Borders, Data With Borders: How Dataspaces Unlock Trusted Cross-Institutional Research Presenter: Kheeran Dharmawardena |
| 13:55–14:10 | Session 11 Question Time & Discussion (15 mins) | Session 12 Question Time & Discussion (15 mins) |
| 14:10–14:45 | Session 11 Breakout & Chat Rooms (35 mins) | Session 12 Breakout & Chat Rooms (35 mins) |
| 14:45–15:00 | Coffee Break / Screen-Off Eye Break (15 mins) | |
| Time | Session 13: Skills, Training & Careers | Session 14: Data Platforms & Data Spaces |
| 15:00–16:55 (115 mins) | Interactive Workshop: Building Diverse and Inclusive Technical Teams: From Recruitment to Retention Presenters: Rowland Mosbergen, Marion Weinzierl (115 mins) | Interactive Workshop: Beyond Interoperability: Hands-On Federated ML for Research Data Curation Infrastructure Presenter: Arnab Mukherjee (15:00–16:30, 90 mins) |
| 16:55–17:00 | Day 3 Closing & Announcements (5 mins) | |
Day 4 – Friday 28 August 2026
AI, Emerging Tools, Reproducibility, FAIR, Open Research, Equity, Access, Knowledge, Sustainability & Impact
Time: 12:30 – 15:30 AEST (UTC+10)
| 12:30–12:40 | Day 4 Welcome & Agenda Briefing (10 mins) | |
| Time | Session 15: AI & Emerging Tools | Session 16: Reproducibility, FAIR & Open Research |
|---|---|---|
| 12:40–13:55 (75 mins) | • 12:40–12:55 (15 mins): AI-driven risk stratification of tumor subpopulations from scRNA-seq using Phenotype Algebra Presenter: Namrata Bhattacharya • 12:55–13:10 (15 mins): Transitioning to Minimally Invasive Diagnostics for Atopic Dermatitis: Validating a Keratin-Based Genetic Signature in Skin Shave Samples Presenter: Rachmat Puziyanto • 13:10–13:25 (15 mins): Transcriptomics and protein-protein interaction analysis to reveal pathogenicity and antimicrobial resistance in children under 5 years with acute diarrhea Presenters: Linda Erlina, Fadilah Fadilah, Asmarinah Asmarinah, Badriul Hegar, Aryo Tedjo • 13:25–13:40 (15 mins): Research Software Without Borders: Building an Open Community Blueprint for Agentic AI in Research Presenter: Trish Radotic | • 12:40–12:55 (15 mins): Open datasets: breaking down the barriers for collaboration and reuse Presenter: Genevieve Buckley • 12:55–13:10 (15 mins): Research Software Without Borders: Open and Reproducible Computational Neuroscience Workflows Across the Globe Presenter: Chitaranjan Mahapatra • 13:10–13:25 (15 mins): From Internal Tool to Research Software: Preparing an R Package for CRAN Presenter: Jyoti Bhogal • 13:25–13:40 (15 mins): Reproducibility Debt: A Hidden Sustainability Challenge in Research Software Presenter: Zara Hassan • 13:40–13:55 (15 mins): Identification of Key Hub Genes in Hepatocellular Carcinoma via Integrated Bioinformatics Analysis and Stringent PPI Network Filtering Presenter: Tazkia Salma Hanifa |
| 13:55–14:10 | Session 15 Question Time & Discussion (15 mins) | Session 16 Question Time & Discussion (15 mins) |
| 14:10–14:45 | Session 15 Breakout & Chat Rooms (35 mins) | Session 16 Breakout & Chat Rooms (35 mins) |
| 14:45–15:00 | Coffee Break / Screen-Off Eye Break (15 mins) | |
| 15:00–15:30 | Day 4 Closing Keynote & RSAA26 Conference Farewell Presenter: Prof. Mohammad Asif Khan (30 mins) | |
Presentation Abstracts and Details
From Method to Research Software, and Beyond: 18 Years of Building mixOmics
- Presenter: Prof. Kim-Anh Lê Cao
- Authors: Prof. Kim-Anh Lê Cao (Melbourne Integrative Genomics, School of Mathematics and Statistics, The University of Melbourne)
Abstract: Biology has become a data-intensive science: a single study now profiles biological samples across hundreds of thousands of molecular features spanning genes, proteins, and other molecules. Integrating these heterogeneous, high-dimensional data into a coherent biological or clinical picture is a hard statistical problem. Statistical methods are only useful once they are packaged into research software that other researchers can run. I have been building the R toolkit mixOmics since 2009, and developing the statistical methods behind it since 2008.
I will share what eighteen years of developing and maintaining mixOmics have taught us. The toolkit is now used by more than 50,000 researchers a year across 155 countries, with 85 packages depending on it. I will describe the architectural decisions that have kept the package extensible, how we preserve backward compatibility as the codebase grows, and what we have learnt from our discussion forum, our workshop programme and user interviews about the researchers we support. I will close on our current challenge: extending mixOmics into a commercial offering, mixOmics Pro, to sustain the open-source research software for the long term.
Presenter Bio: Kim-Anh Lê Cao is Professor of Statistical Genomics at the University of Melbourne and Director of Melbourne Integrative Genomics. She has secured three consecutive NHMRC fellowships since 2014, and numerous awards for her contributions to statistics applied to molecular biology, including the Moran medal from the Australian Academy of Science. Her research focuses on multi-omics data integration and she leads the development of the popular R toolkit mixOmics, now extended into the spin-out mixOmics Pro.
An Efficient High Dimensional Gene Expression Based Multi-Class Cancer Prediction Model Using Ensemble Learning
- Presenter: P. Krithy Sreshta
- Authors: P. Krithy Sreshta (SRM Institute of Science and Technology)*
Abstract: Cancer continues to be a predominant worldwide health concern, and early, non-invasive identification significantly improves survival rates and treatment results. Conventional diagnostic techniques, including tissue biopsies and imaging, are often invasive, costly, and time-consuming. Recent improvements in next-generation sequencing facilitate the examination of gene expression data at the molecular level. Nevertheless, this data is high-dimensional and noisy, making manual analysis impossible. This paper presents a machine learning framework for multi-class cancer categorization using Tumor-Educated Platelet (TEP) gene expression data. The dataset undergoes preprocessing using normalization, feature selection, and Principal Component Analysis (PCA) to reduce data dimensionality. Several supervised models, including Support Vector Machine (SVM), Random Forest, Logistic Regression, and XGBoost, were trained and refined by cross-validation. The performance is assessed using accuracy, precision, recall, F1-score, and ROC-AUC metrics. Experimental findings demonstrate that ensemble models, notably Random Forest and XGBoost, surpass other classifiers in managing high-dimensional biological data. The finished model is then deployed via a web-based application for real-time, user-friendly cancer prediction. This study presents a scalable and non-invasive method for precise cancer categorization that facilitates future clinical decision-making systems.
Presenter Bio: I am a Post graduate student at SRM Institute of Science and Technology with a strong interest in Artificial Intelligence, Machine Learning, and Data Analytics. My research focuses on applying advanced machine learning techniques to healthcare data, particularly in cancer prediction using gene expression analysis. I have worked on developing efficient multi-class classification models using ensemble learning methods to improve prediction accuracy. I am passionate about leveraging AI for real-world problem solving and contributing to innovations in medical data analysis. I have experience in Python, data preprocessing, and model evaluation techniques. My work aims to bridge the gap between computational intelligence and healthcare applications, making predictive systems more reliable and impactful.
The RCP Discovery Internship Program: Building a Global Innovation Lab for Research Software
- Presenter: Rowland Mosbergen
- Authors: Rowland Mosbergen (WEHI)*
Abstract: The RCP Discovery Internship Program at WEHI (Walter and Eliza Hall Institute of Medical Research) presents a novel model for advancing research software through a globally distributed, fully online internship framework. Since its launch in July 2021, the program has engaged over 300 interns from six continents across 14 intakes, while proving that research software complexity can be addressed efficiently with minimal resources while reducing risk and maximising efficiency and effectiveness. This presentation describes how the program functions as an innovation lab for interns: recruiting intentionally multi-disciplinary teams, embedding DEI principles, and deploying a “best effort” mentorship philosophy that prioritises critical thinking, adaptability, and tolerance for ambiguity over rigid deliverables. We outline the operational strategies that enable a single supervisor to manage cohorts of up to 40 interns simultaneously, including structured documentation practices, Individual Learning Plans, Technical Diaries, asynchronous feedback mechanisms, with a focus on guided autonomy. We also showcase how this program has innovated for researchers, highlighting three key outcomes: Bioconductor package compliance, streamlined flow cytometry workflows via automated file merging, and the REDMANE research data ecosystem that is now available as a demo. We also highlight how this program has created a new effective methodology for tackling complex software projects that moves from proof-of-concept through to handover, deliberately surfacing complexity early rather than deferring it.
Presenter Bio: Rowland is a strategic leader in digital research infrastructure with 25+ years’ experience translating HPC, AI, data and cloud technologies into scalable research and industry services used by global communities. Known for building partnerships across researchers, national infrastructure providers and technology teams to deliver high-impact digital capability, Rowland provides strategic direction, inclusive leadership and scalable programs that accelerate research adoption and demonstrate measurable impact. Rowland co-founded RSAA, RSAfrica, and RS Latin America, and established Equersa, the umbrella organisation supporting all three regional conferences. He also serves on the steering committee of the RSE Australia New Zealand Association.
Building Capacity from the Inside: A Project Update from the Research Software Engineering Capacity Enhancement Project (RSE-CEP)
- Presenter: James Smithies
- Authors: James Smithies (The Australian National University)*
Abstract: The Research Software Engineering Capacity Enhancement Project (RSE-CEP) is a co-investment partnership with the Australian Research Data Commons (ARDC) through the Humanities and Social Sciences (HASS) and Indigenous Research Data Commons (DOI: 10.3565/x5f6-mw53). The project is documenting software patterns, a technology roadmap, and an architecture for the ARDC Community Data Lab (CDL), a deliverable of the Humanities and Social Sciences and Indigenous Research Data Commons (HASS-I RDC). We aim to support the development of Australia’s HASS research infrastructure and technical community (RI) using a collaborative approach that includes consultation with key HASS-I RDC projects. This talk will provide an update on the project’s progress, after ~12 months of operation. Significant conceptual work has been undertaken to define the nature of HASS software patterns, and to understand how those patterns are best used to design the supporting roadmap and architecture. We’ll share emerging insights into how our findings work together to guide decision-making over time and we’ll show how the project’s outputs will contribute to the technical development of Australia’s digital HASS RI, by exposing quality patterns and practices, and demonstrating the connection between technical and strategic development. It is becoming clear that projects like RSE-CEP can increase the potential for equity, diversity, and sustainability by building shared understanding across the RSE and funding communities, helping us shape the future of HASS RSE from the inside.
Presenter Bio: James started working at the Australian National University in 2024, as founding director of the HASS Digital Research Hub. Before ANU, he was Professor of Digital Humanities in the Department of Digital Humanities at King’s College London and founding director of King’s Digital Lab. He also served as deputy director of eResearch at King’s. James has also worked in the government and commercial IT sectors in the United Kingdom and New Zealand.
Launching Discworld: A bespoke, user-considerate, CryoEM processing cluster
- Presenters: Nathan Glades, Keiran Rowell
- Authors: Keiran Rowell (Structural Biology Facility - UNSW)*; nathan Glades (Structural Biology Facility - UNSW)
Abstract: CryoEM is one scientific domain where skilled domain experts, instrument scientists, and biosciences higher-degree research candidates encounter the demands of high performance computing workloads. The work is visual, both in graphical processing and expert graphical interpretation. The software written for it often assumes the resources of a dedicated workstation, which works as edge compute for responsive analysis but doesn’t scale to multi-user or long timeframe workloads. We approached the problem to provide workstation-like accessible features, while allowing the researcher to leverage the mental model and power of 24/7 scheduled system. There are many approaches to this cross-disciplinary computing demand, and we’re presenting some choices in the making of a small bespoke cluster designed and administered to lower both researcher friction and the software & hardware administration burden. This involves a web portal for browser desktops via noVNC. Its access and networking implications will be discussed. Modern Rust CLI toolchains with researcher considerate behavior are provided by default, allowing performant searching work on the massive file burden of CryoEM data. There is a proliferation of heterogeneous processing software –often with competing resource access designs– so it will be discussed how we expose these pre-existing software package management and module systems to the researcher while providing HPC support. This also involves co-tenancy of various computational methods and web applications under the same scheduler, so the fairshare algorithm can efficiently apportion resources. Training researchers with accessible ‘mental model’ computing demos & documentation to become exploratory and self-sufficient is also a focus, so they can spend more time learning and practicing their scientific expertise.
Presenter Bio: Nathan: Nathan is a systems administrator with a background in Linux and High Performance Compute. He currently works at the Structural Biology Facility in UNSW, where he builds and administers high performance compute systems for bespoke bioscience applications. His role involves training and assisting users with their computational demands, architecting systems to meet such demands, and administering and maintaining this infrastructure. Currently, he is focused on meeting the demands of CryoEM workflows, and building a scalable and secure system for such workflows. Nathan is driven by solving tricky and bespoke architectural and systematic computational problems. He also strives to deliver access models that allow people to work smarter, not harder. Keiran: Keiran is a computational scientist, with a strong background and deep enthusiasm for molecules and the fascinating things that they do. He currently works at the Structural Biology Facility at UNSW, where he runs calculations.
Open datasets: breaking down the barriers for collaboration and reuse
- Presenter: Genevieve Buckley
- Authors: Genevieve Buckley (Walter and Eliza Hall Institute)*
Abstract: Open data is an invaluable tool for the research community, inviting open collaboration and reuse. The availability of open datasets promotes FAIR standards, meaning that data is findable, accessible, interoperable, and reusable. This talk will provide an overview of a variety of freely available open datasets. They cover a wide range of data types and areas of research, including genomics, health, neuroscience, imaging, social science and more. We will explore some case studies of software tools made to support access and allow more informative use of existing open datasets. We will discuss how you can integrate open data in combination with your existing research projects. Researchers can cross reference to a common standard, allowing more accurate comparisons between new findings. Open data can provide wider context for a specific research study. Most importantly, they can be hugely useful for data mining in for entirely new research projects. This is a fantastic way for the community as a whole to benefit from existing, previously funded research. Finally, we will demonstrate to researchers the benefits of contributing your own research data and software. The benefits are not simply altruistic. This talk will explain some common ways that people typically choose to do this, including suitable locations for hosting, creating digital object identifiers.
Presenter Bio: Genevieve Buckley is a research software engineer at the Walter and Eliza Hall Institute of Medical Research. She builds software tools to solve problems in scientific discovery. Her interests include deep learning, automated analysis, and contributing to open-source software projects.
AI-driven risk stratification of tumor subpopulations from scRNA-seq using Phenotype Algebra
- Presenter: Namrata Bhattacharya
- Authors: Namrata Bhattacharya (Peter MacCallum Cancer Centre)*
Abstract: Robust characterization of cellular phenotypes from single-cell gene expression data is of paramount importance in studying complex biological systems and diseases. Single-cell RNA-sequencing (scRNA-seq), coupled with robust computational analysis, facilitates the characterization of phenotypic heterogeneity in tumors. However, given the extent of intra-tumoral heterogeneity, assessing the risk associated with individual malignant cell subpopulations is challenging, primarily due to the complexity of the cancer phenotype space and the lack of clinical annotations associated with tumor scRNA-seq studies. To address this challenge, we have proposed an unsupervised transfer learning-based method called SCellBOW, which facilitates the identification and risk stratification of tumor subpopulations using language modeling techniques. SCellBOW can precisely identify phenotypically divergent cell types from tumor scRNA-seq gene expression profiles. Furthermore, SCellBOW offers a remarkable feature called phenotype algebra, which enables effective risk stratification of individual malignant cell subpopulations by executing algebraic operations such as ‘+’ and ‘–’ on the single-cells in latent space. This feature allows for the simulation of the residual phenotype of tumors after the positive and negative selection of specific malignant cell subtypes in a tumor. SCellBOW has proven effective in characterizing poorly defined metastatic prostate cancer. Using SCellBOW, we identified a hitherto unknown subpopulation of de-differentiated cells in metastatic prostate cancer, AR−/NElow (androgen receptor negative, neuroendocrine-low), that exhibit greater aggressiveness and have distinct gene expression patterns not observed in any previously identified molecular subtypes. The risk-stratification capabilities of SCellBOW hold promise for tailored therapeutic interventions by identifying clinically relevant subpopulations and understanding their impact on prognosis.
Presenter Bio: Dr Bhattacharya is a Postdoctoral at the Peter MacCallum Cancer Centre, working with clinician-scientist Prof. Mark Dawson. She specialises in computational cancer genomics and artificial intelligence (AI) applications in oncology. Her work focuses on developing and applying advanced bioinformatics and AI frameworks for tumour microenvironment deconvolution, therapeutic target discovery and response prediction, with an emphasis on translational cancer research.
Namo: Working with Named Dimensional Data
- Presenter: thoran
- Authors: thoran - (NA)*
Abstract: Research data arrives as records, which are rows with named fields from databases, CSV files, and APIs. Yet the dominant libraries for processing data require researchers to reshape that data into positional, numeric structures before they can work with it. Names can be lost. Indices become opaque. Boilerplate accumulates.
Namo is a library for named dimensional data that works with row-based data directly. It takes arrays of hashes and makes dimensions first-class:
data[gene: 'BRCA1', tissue: 'breast']
selects by name;
data[:category] = proc{|row| row[:expression] > 20 ? 'high' : 'low'}
defines a computed dimension that carries through filtering and composition;
data * metadata
joins two datasets on shared dimensions with a single operator.
Set operations (|, &, -), projection, contraction, and Ruby’s built-in collection methods all work on every Namo out of the box.
The talk introduces Namo’s API by describing each of selection, projection, contraction, formulae, chaining, and composition, then through a sequence of live examples in irb (interactive Ruby), demonstrates each in practice.
This is followed by a comparison of the same operations in pandas, xarray, and R. It also discusses how the library was designed and built in conjunction with Claude Code.
Finally, it presents proposed future syntax changes: from the current hash-based row access (1.x) to bare names (2.x) to a declarative define-block DSL (3.x), with each step producing simpler, more readable code while preserving meaning.
Presenter Bio: thoran is an independent software developer based in Melbourne, Australia, with over 40 years of software experience. He maintains over 20 published Ruby libraries spanning areas from HTTP, IMAP, and WebDAV clients, to state machines, utility classes, and API clients. His work sits at the intersection of software tooling, data infrastructure, and systems design. Some of his current projects include Namo (named dimensional data in Ruby), Configured (declarative cross-platform machine configuration and deployment), and Cryptarium (sovereign encrypted data containers). He is a long-standing advocate for Ruby’s applicability beyond web development. His approach to software development emphasises compositional design, minimal dependencies, and interfaces that feel inevitable rather than learned.
Breaking Barriers in Research Data: Implementing FAIR Workflows and Persistent Identifiers Across Distributed Systems
- Presenter: Aline de Souza Andrade
- Authors: Aline de Souza Andrade (UWA)*
Abstract: Research outputs are often distributed across multiple institutions, repositories, and disciplines, creating significant challenges for discoverability, interoperability, and reuse. These challenges are not only technical but also organisational, limiting the ability of researchers and stakeholders to effectively access and connect research outputs. This presentation shares practical insights from a cross-institutional initiative within the Australian National Environmental Science Program (NESP) Resilient Landscapes Hub, focused on improving research data practices through the implementation of FAIR principles and Persistent Identifiers (PIDs). The project was designed to address inconsistencies in metadata, limited use of identifiers, and fragmentation across repository systems. A key component of the work involved engaging with multiple institutions and repository platforms to identify both technical limitations and social barriers to adoption. Challenges included differences in repository capabilities, lack of standardisation in metadata fields, and varying levels of awareness and engagement among researchers and repository managers. To address these issues, the project developed a practical Data Model Guideline to support researchers in creating consistent, high-quality metadata and implementing PIDs such as DOIs, ORCID, ROR, and the Research Activity Identifier (RAiD). The approach was tested through real-world implementation across several institutions, including the integration and harvesting of metadata into Research Data Australia (RDA). The presentation will highlight how connecting projects, people, organisations, and outputs through structured metadata and identifiers can significantly improve discoverability and interoperability. It will also reflect on key lessons learned, particularly the importance of addressing social challenges alongside technical solutions when building sustainable, cross-institutional research data ecosystems. This work demonstrates how practical, community-driven approaches can help break down institutional barriers and contribute to more connected and reusable research systems.
Presenter Bio: Aline de Souza Andrade is a Project Officer at The University of Western Australia, working with the NESP Resilient Landscapes Hub. Her work focuses on improving research data management practices through the implementation of FAIR principles, metadata standards, and persistent identifiers. She has led cross-institutional initiatives in collaboration with the Australian Research Data Commons (ARDC), supporting the integration of research outputs into national infrastructure such as Research Data Australia (RDA). Aline has worked closely with researchers and repository managers across multiple organisations to address technical and social challenges in metadata standardisation and interoperability. Her work contributes to building more connected, discoverable, and reusable research data ecosystems.
Research Software Without Borders: Open and Reproducible Computational Neuroscience Workflows Across the Globe
- Presenter: Chitaranjan Mahapatra
- Authors: Chitaranjan Mahapatra (Department of Electronics and Communication Engineering SRM Institute of Science and Technology, Delhi NCR Campus, India)*
Abstract: Research software is increasingly central to modern computational biology, neuroscience, and biomedical engineering, yet access to sustainable software ecosystems remains uneven across regions and institutions. This presentation shares practical experiences from developing and applying open, reproducible, and cross-platform computational neuroscience workflows across the globe with international collaborations. Drawing from research in electrophysiology, ion-channel modeling, computational neuroscience, and AI-enabled biomedical systems, the talk will demonstrate how open-source scientific software, interoperable modeling frameworks, and community-driven practices can reduce barriers for researchers working in resource-constrained environments. The presentation will highlight the use of Python-based scientific computing ecosystems, simulator-independent neuronal modeling tools such as PyNN, reproducible workflows, and collaborative research software practices developed through international projects including the Human Brain Project and global computational neuroscience communities. The talk will also reflect on challenges faced by researchers from the Global South, including limited computational infrastructure, fragmented training opportunities, language and accessibility barriers, and sustainability issues in long-term software maintenance. Examples from mentoring, open tutorials, conference workshops, and community-led training initiatives will illustrate how collaborative and inclusive research software communities can improve scientific participation and reproducibility. By connecting experiences from academia, international collaborations, and interdisciplinary biomedical research, this presentation aligns with the RSAA26 theme of “Research Software Without Borders” and emphasizes equitable access, open science, and sustainable research software practices that support diverse global communities.
Presenter Bio: Chitaranjan Mahapatra is a visiting postdoctoral researcher at the Institute for Basic Science, Daejeon, South Korea. He is joining as a contractual research faculty at SRMIST, Delhi, India, from July 2026. His research spans computational neuroscience, electrophysiology, biomedical engineering, AI-enabled healthcare systems, and open scientific computing. He has worked with the Human Brain Project and open computational neuroscience communities. He has organised tutorials on PyNN and computational modelling, served as reviewer and programme committee member for multiple international scientific conferences, and mentored students through initiatives such as Neuromatch Academy and Google Summer of Code-related activities. His work emphasises reproducible computational workflows, interdisciplinary research software, and accessible scientific training for researchers across diverse institutional and geographic settings.
CPUX: State-Driven Orchestration for Distributed Research Simulations
- Presenter: Pronab Pal
- Authors: PRONAB PAL (KEYBYYESYSTEMS)*; Pronab Pal (Keybytesystems)
Abstract: Scientific workflows are commonly orchestrated through explicitly authored dependency graphs, even when the simulation stages themselves already declare what data they require and what they produce. This creates maintenance overhead, reduces composability, and becomes especially awkward in distributed settings where data locality and sovereignty matter. This presentation introduces CPUX (Common Path of Understanding and Execution), a state-driven orchestration model in which execution order emerges from typed shared state rather than from a separately maintained workflow graph. In CPUX, each simulation stage is represented as a Design Node inside an Intention Container, declaring only the Pulses it requires and the Pulses it emits. A readiness procedure, synctest, checks whether the required Pulses are present in the shared Field. If they are present, the stage runs; if not, it waits. As output Pulses accumulate, downstream execution order emerges automatically. Using a working Go-based prototype, I will present a three-stage distributed simulation pipeline—population, traffic, and hospital-capacity planning—in which the observed execution sequence emerges without explicit dependency wiring. I will also show how CPUX supports geographically distributed execution by allowing containers to run on region-appropriate hosts while sharing only derived summary state. The core claim is that CPUX replaces explicit control flow with a fixed-point computation over shared semantic state, offering a practical model for modular, reproducible, and sovereignty-aware research software.
Presenter Bio: Pronab Pal is a Systems Architect and the founder of Keybyte Systems, based in Melbourne, Australia. With a career spanning over two decades in cloud infrastructure, mobile systems, and distributed architectures, he approaches research software from the perspective of an engineer optimized for real-world scaling and cross-platform resilience. His recent work focuses on “Intention Space”—a cognitive-architectural layer that replaces explicit control flow with state-driven, fixed-point orchestration (CPUX). By developing software that aligns machine execution with situational awareness and human intention, Pronab seeks to create more observable, reproducible, and “perceptive” computational systems. He holds an alumnus status from the Indian Institute of Science and remains an active contributor to the independent research community through Intentix Lab.
Intelligent Automation in Healthcare Innovation
- Presenter: Anushka Bhargava
- Authors: Anushka Bhargava (Philips)*
Abstract: Healthcare applications of artificial intelligence are becoming increasingly sophisticated, moving away from the limitations of standard machine learning approaches to more advanced systems that reason, plan, and autonomously perform complex tasks. This includes the use of intelligent AI agents in applications as diverse as medical imaging tools and platforms, research databases and data analysis, and even workflow orchestration systems. In this lightning talk, we will cover some of the key trends around the application of agentic artificial intelligence and automation within healthcare ecosystems, exploring how such innovations can increase interoperation, enable streamlined research workflows, and mitigate institutional friction. Through specific examples of AI technology in use in actual healthcare software environments, including applications for imaging platforms and automated data processing pipelines, among other examples, this talk will explore the future of agentic AI within the space of healthcare research software. Finally, we will explore the potential role of research software communities in developing sustainable, equitable AI technologies for healthcare.
Presenter Bio: Anushka Bhargava is currently a software development engineer at Philips Healthcare, focusing on software platforms for healthcare, medical image processing applications, workflows, and automation systems with AI integration. Some of her current projects include integrating MCPs with Copilot for automation purposes, optimising clinical visualisation tools such as MRViewer and WPF toolkit, and developing workflow orchestration and data migration systems. Previously, she was an intern software development engineer at Amazon, where she developed scalable backend systems for high-transaction systems. In her previous roles, Anushka has worked with AI and its applications, particularly in the healthcare sector, including deep learning, machine learning, medical data processing, and intelligent automation systems. She has also been involved in research projects at several institutions, including the DRDO and the National University of Singapore.
Ecosystem Indicators Workflows: Building a FAIR, Modular Toolbox for Reproducible Ecological Assessment in Australia
- Presenters: Abhimanyu Raj Singh, Jose Ferrer Paris
- Authors: Abhimanyu Raj Singh (UNSW)*; Jose Ferrer Paris (UNSW)
Abstract: Ecosystem-specific indicators underpin ecosystem accounting, risk assessment, and environmental reporting yet their development across Australia remains fragmented. Methods are often ad-hoc, less documented, and difficult to scale across the country’s diverse ecosystems. This creates a reproducibility gap that limits scientific rigour and policy uptake. The Ecosystem Indicators Workflows project (a co-investment partnership with the ARDC, February 2026–2028) is building a modular toolbox of transparent, reproducible workflows to guide users through indicator development from raw data to analysis-ready outputs. These workflows use conceptual models to ensure alignment of ecosystem knowledge with the intended purpose and are designed with FAIR principles embedded from the outset. Particularly we aim to (a) enable integration of existing ecosystem distribution data and cross-referencing national and global ecosystem classifications, (b) adopt standard controlled vocabularies to improve the findability and semantic interoperability, and (c) use persistent identifiers and a publication process to promote reusability in the research community. In this lightning talk, I share early lessons from requirements gathering with ecologists, government agencies, and conservation managers across six institutions. Key challenges identified in this first stage of the project include the need to accommodate different user expertise levels, the variety of data types, and the accessibility of relevant data sources. The architectural choices required to balance scientific rigour with accessibility, including how we are approaching workflow modularity, and semantic alignment will be discussed. This work-in-progress invites feedback from RSEs who have navigated similar trade-offs in reproducible research infrastructure.
Presenter Bio: Abhimanyu Raj Singh is a Data Ecologist with the Ecosystem Indicators Workflows (EIW) project at UNSW Sydney and EcoCommons/QCIF. He delivers workflow packages, training materials, and community engagement activities, bridging ecological science and data infrastructure. Abhimanyu is committed to open science and building technical capacity across government, academic, and industry sectors engaged in ecosystem monitoring and reporting. Dr Jose Ferrer Paris develops methods and frameworks for ecosystem and biodiversity conservation, influencing global and national policy. His research applies spatial ecology and biodiversity informatics to species distributions, camera trap data, species traits, and spatio-temporal ecosystem indicators feeding digital knowledge bases for policy, risk management, restoration, and natural capital accounting. Committed to reproducibility and open science, he shares code and data openly, supporting evidence-based decisions across diverse user communities.
A 3-month challenge to revive, share, and sustain a PhD-era bioinformatics tool
- Presenter: Saliha Elif Yildizhan
- Authors: Saliha Yildizhan (Peter MacCallum Cancer Centre)*
Abstract: During my PhD, I developed a Shiny application for classifying pancreatic cancer subtypes using pairwise gene expression comparisons (https://salihaey.shinyapps.io/pcsubtypes/). The tool underpinned a preprint that was never formally published (https://www.biorxiv.org/content/10.1101/2025.01.13.632859v1), yet the app itself is still running, still accessible, and still waiting to be found. Over the three months leading up to RSAA26, I will attempt to turn this dormant side-project into a robust and visible research tool. Through this process, I aim to increase knowledge across the following areas: making research software discoverable beyond its immediate context; large-scale data analysis platforms; and using AI for tool development and documentation, reflecting honestly on how useful it is.
Presenter Bio: Saliha holds a bachelor’s degree in Molecular Biology and Genetics from Boğaziçi University and a PhD in Biostatistics and Medical Informatics from Acıbadem University. Her doctoral work involved developing new methods for biomarker discovery that are sensitive to feature interactions in machine learning models, aiming to provide mechanistic insights into neuron–tumour cell interactions in pancreatic cancer. Since joining the Mark Dawson Lab in 2024, her research has focused on developing novel computational approaches to uncover the biological mechanisms underlying transcriptional plasticity in therapeutic resistance.
Research Software Engineering in the Age of Generative AI
- Presenters: Richard Littauer, Michelle Barker, Sandra Gesing
- Authors: Richard Littauer (Te Herenga Waka Victoria University of Wellington)*; Michelle Barker (ReSA); Sandra Gesing (US-RSE)
Abstract: The Research Software Engineering in the Age of Generative AI: Building a Community Vision workshop, held in March in Edinburgh, UK, brought together participants to explore how Generative AI may reshape the research software ecosystem, and to help inform a broader community vision for the future of the field. The workshop addressed both the opportunities and risks emerging from the increasing integration of AI into scientific workflows, including faster software development, automation, and broader accessibility, alongside concerns around reliability, reproducibility, provenance, verification, security, and workforce transformation. Through lightning talks, community discussions, surveys, and focused working groups, participants examined how RSE roles, institutional policies, collaboration models, training approaches, and software verification practices may evolve in the coming years. Discussions highlighted the continuing importance of human expertise, trustworthy software practices, and coordinated community action in an AI-enabled research ecosystem. A major outcome of the workshop was the identification of more than 50 proposed pilot activities and studies spanning nine thematic areas: institutional policy and narratives, tradeoffs and risk frameworks, software attribution and incentives, verification and validation, evolving RSE roles, training and workforce development, management playbooks, equitable access to AI technologies, and new forms of collaboration between researchers and AI systems. Proposed activities range from lightweight community-driven initiatives and training communities of practice to large-scale empirical studies, ethnographic research, and longitudinal analyses of AI-assisted research software development. This talk will give an overview of the pilots proposed.
Presenter Bio: Richard Littauer is an organizer for CURIOSS, the Community for University and Research Institution Open Source Program Offices (OSPOs), and SustainOSS. He is also a PhD student in Computer Science at Te Herenga Waka Victoria University of Wellington in Pōneke, Aotearoa New Zealand. His research interests beyond involve open science, open source, community science platforms, and taxonomy.
The Future (and Present) of Research Software Engineering with Cross-functional Teams
- Presenter: Gregory B. Poole
- Authors: Gregory Poole (Swinburne University of Technology)*
Abstract: I’m aiming for this to be an informal discussion about the current status and future of research software engineering and the place of cross-functional teams (if you don’t know what those are, I’ll explain) within that. I’ll come armed with my experiences from a decade of working with Astronomy Data and Computing Services (ADACS; again: if you don’t know what that is, I’ll explain) and lots of questions FOR YOU about the challenges you are facing, the things that are holding you back, etc.. So please come armed with thoughts about what - in terms of research software - is going well for you, what’s not and where you think things are headed.
Presenter Bio: Gregory Poole is the Astronomy Data Science Coordinator for Swinburne University. In this role he manages the activities of ~15 research software engineers and software professionals contributing to the NCRIS-funded Astronomy Data and Computing Services (ADACS) program. Previously he has been a postdoctoral researcher at the University of Melbourne and Swinburne University where he specialised in using distributed computing systems to generate and analyse large cosmological simulations of the Universe. He obtained his PhD studying hydrodynamical simulations of merging galaxy clusters at the University of Victoria, Canada.
Fortran is cool don’t let your friends convince you otherwise
- Presenter: Jorge Luis Galvez Vallejo
- Authors: Jorge Galvez Vallejo (NCI)*
Abstract: Fortran has been around for almost 65 years, and while C/C++ and now Rust have taken software engineering careers by storm, Fortran still reigns supreme in climate, weather, and computational fluid dynamics. Most of these codebases, however, have been driven by scientists rather than software engineers — and the result is a mountain of technical debt that makes maintenance and innovation harder than it should be. This is not a failing of the scientists; it’s an opportunity for the research software engineering community. In this talk I will argue that Fortran is a modern programming language worth taking seriously, and that the RSE community has a real role to play in shaping its future. Modern Fortran offers object orientation, (almost) safe dynamic memory handling, arrays as first-class citizens, and native parallelism baked into the language. Paired with the right compiler and good practices, the same code can run on a single core, across many cores, and on GPUs with minimal changes — a portability story that is hard to match in any other language. And critically, Fortran lets domain scientists write fast code without first having to become systems programmers. I will close with a tour of the modern Fortran ecosystem: the Fortran Package Manager, the Fortran Standard Library, and the pic series of libraries I have been developing to provide portable, standard-library-level routines in native Fortran.
Presenter Bio: Jorge is a computational scientist with experience in computational chemistry and recently in geophysical fluid dynamics. His main interest is performance portability of GPU programming models across programming languages, and enabling access of massively parallel architectures. He used to do a lot of C++ but now prefers Fortran as his main programming language. Jorge is obsessed with CMake and unit testing and will implement these in your codebase if you’re not careful enough.
BRI DataLab: an AI-Assisted Research Platform for Infrastructure Finance Research in International Relations
- Presenter: Sanoop Sajan Koshy
- Authors: Sanoop Sajan Koshy (Indian Institute of Technology Madras)*
Abstract: Research on China’s Belt and Road Initiative (BRI) is often fragmented across datasets, policy documents, institutional reports, and academic literature. Researchers working on infrastructure finance and international political economy frequently navigate disconnected sources, making integrated analysis time-intensive and difficult to reproduce. This presentation introduces BRI DataLab, an open-source AI-assisted research platform developed to support research on Chinese infrastructure finance. The platform combines a harmonized infrastructure finance dataset derived from AidData’s Chinese development finance records, a curated corpus of policy and academic documents, and a retrieval-augmented AI interface designed for transparent querying and synthesis. Rather than presenting AI as a replacement for scholarly expertise, the project explores how AI-assisted workflows can support research practice in data-fragmented domains. The presentation will demonstrate how the platform enables researchers to connect structured project-level data with unstructured textual sources within a single research environment. The session will focus on practical lessons learned from building the system as an interdisciplinary social science project with limited resources, including challenges related to data harmonization, document retrieval, transparency, prompt design, and methodological constraints. It will also reflect on broader questions surrounding reproducibility, open research infrastructure, and responsible AI-assisted workflows in the social sciences.
Presenter Bio: Sanoop Sajan Koshy is a PhD Candidate in International Relations at the Indian Institute of Technology Madras. He was also a recipient of the Chinese Government Scholarship (2024-25) as an exchange scholar at the Beijing Normal University. His research focuses on China’s Belt and Road Initiative, infrastructure finance, and the political economy of development. His recent work examines the intersection of artificial intelligence, digital research methods, and computational social science.
Computational Research Sustainability
- Presenter: Pao Corrales
- Authors: Paola Corrales (21st Century Weather Centre)*
Abstract: As research becomes increasingly data- and compute-intensive, the environmental impact of computational research is becoming harder to ignore. Every computation activity, in my case, processing large datasets and running climate and weather models, consumes energy and contributes to carbon emissions. At the same time, these activities are essential. Computational research across fields such as climate science, health, environment, and many other areas plays a critical role in addressing major global challenges. At the same time, it is equally important to acknowledge their cost and to ensure we are using our computational resources as efficiently and responsibly as possible. This mini workshop invites research software engineers, researchers and anyone interested to join a conversation about sustainability in computational research. We’ll explore what green computing means in practice, how different initiatives and tools can help quantify and reduce the digital footprint of research, and how these ideas can be adapted within our own institutions and infrastructures. The session will highlight existing tools, initiatives, and community efforts aimed at measuring and reducing the carbon footprint of research. Participants will also discuss practical actions that can be adopted within their own projects and institutions, from improving software efficiency and workflow design to increasing awareness around sustainable computing practices. By sharing experiences across disciplines and institutions, this workshop aims to foster a broader conversation around sustainability in research software and help build a community of practice where efficiency, transparency, and environmental responsibility are considered part of good research and software engineering.
Presenter Bio: Pao has a PhD in Atmospheric Science from the University of Buenos Aires. During her PhD, Paola applied data assimilation techniques to improve the representation of mesoscale convective systems and associated precipitation. In particular, her research focused on data assimilation of conventional observations and radiances from polar and geostationary satellites. She has experience working with Numerical Weather Prediction models and developing software using R and Fortran. She is also part of R-Ladies, rOpenSci and The Carpentries and is interested in open source, open science and reproducibility. She recently joined the ARC Centre of Excellence for the Weather of the 21st Century as part of the Research Software Engineer Team in Australia, where she supports researchers studying what the weather will look like as our planet keeps warming up.
Software Without Borders, Data With Borders: How Dataspaces Unlock Trusted Cross-Institutional Research
- Presenter: Kheeran Dharmawardena
- Authors: Kheeran Dharmawardena (Cytrax Consulting, Melbourne, Australia)*
Abstract: This conference carries a bold promise: “Research Software Without Borders.” Yet the single biggest barrier to cross-institutional collaboration isn’t software at all - it’s the organisational and institutional arrangements governing data access. The digital society we live in today relies on effective access to data. Data has become one of the most valuable assets on Earth, and in an AI-accelerated world that consumes data voraciously, individuals and organisations are increasingly less willing to surrender control of their data to someone else’s platform. To grant access to data, they need to trust it will be handled on their terms. Today, this trust is achieved through bilateral and multilateral agreements that take months or years to negotiate. While this served us well in the past, it doesn’t scale to meet today’s needs. Europe has spent the last decade solving this problem. They’ve built a framework and architecture for establishing trustworthy, federated data access networks that are now being adopted globally - Dataspace. A dataspace is a third way between the two approaches we are familiar with: centralised platforms where the platform owner holds the power, and point-to-point agreements that don’t scale. In a dataspace, each participant retains full control of their own data - it stays with its owner - but can make it available to others through connectors that enforce machine-readable data contracts under a common trust framework. No single entity owns or controls the ecosystem. This talk will cover why federated data access networks are essential in the data-driven digital society, how the International Data Spaces (IDS) framework provides a blueprint for building such networks, and why we should adopt this approach.
Presenter Bio: Kheeran Dharmawardena is Director and Principal Consultant at Cytrax, a specialist practice focused on the human, organisational, and institutional dimensions of complex data infrastructure initiatives. With 25+ years spanning research, government, and higher education, he specialises in multi-stakeholder programs where success depends on aligning people, policy, and technology. Kheeran led the Australian Dataspaces Program at the Australian Research Data Commons (ARDC), initiating pilot projects across biosecurity, environmental assessment, and legal applicability. He also led the work for the Commonwealth Department of Climate Change to design national governance frameworks and data standards that united all states and territories around a shared approach to cross-jurisdictional biodiversity data access. Kheeran co-chairs interest groups at the Research Data Alliance and has extensive experience building research data infrastructure.
An Integrative Bioinformatics Framework for Identifying Core Regulatory Pathways in Melanoma Metastasis
- Presenters: Tifany Nabilah, Linda Erlina
- Authors: Tifany Nabilah (Master’s Programme in Biomedical Sciences, Faculty of Medicine, Universitas Indonesia)*; Linda Erlina (Bioinformatics Core Facilities, Indonesian Medical Education and Research Institute (IMERI), Faculty of Medicine, Universitas Indonesia)
Abstract: Melanoma metastasis remains the primary driver of patient mortality, yet the molecular mechanisms underlying the transition from primary to metastatic states remain incompletely characterized. Current computational approaches often analyze single datasets, limiting reproducibility and generalizability across patient cohorts. This study implements an integrative bioinformatics workflow to systematically identify core gene networks mediating melanoma progression through rigorous cross-dataset validation. Three independent Gene Expression Omnibus public datasets (GSE7553, GSE46517, GSE15605; total 125 primary melanoma, 91 metastatic melanoma skin biopsies) were evaluated for differentially expressed genes using GEO2R with stringent cutoffs (adjusted p < 0.05, |log₂FC| > 1), identifying 156 intersecting differentially expressed genes. Functional enrichment analysis via ShinyGO v0.85.1 identified six KEGG pathways significantly enriched (false discovery rate < 0.05), predominantly involving in extracellular matrix-receptor interaction (fold enrichment = 8.48), focal adhesion (4.49), and PI3K-Akt signaling (3.75). Protein-protein interaction network construction using STRING database v2.2.0 in Cytoscape v3.10.4 identified 16 nodes with 33 edges. CytoHubba plugin ranked by Maximal Clique Centrality identified ten hub genes (EGFR, SERPINB5, SFN, LAMB3, LAMC2, LAMA3, JUP, CALML5, FGFR2, EFNA3), all consistently downregulated in metastatic samples. To ensure clinical relevance, validation using TCGA-SKCM data (n = 461) demonstrated that four hub genes (SERPINB5, SFN, EFNA3, LAMC2) significantly stratified overall survival (hazard rate 1.4–1.5, log-rank p < 0.022). Notably, stage-specific analysis identified a critical expression shift during the Stage II invasion phase, suggesting these genes as early biomarkers for metastatic transition. This reproducible computational framework unravels melanoma metastasis as systematic dismantling of structural and epithelial regulatory pathways, providing validated biomarker candidates and demonstrating the power of multi-dataset integration in cancer bioinformatics research.
Presenter Bio: Tifany Nabilah is a graduate student in the Master program of Biomedical Sciences, Bioinformatics concentration at Universitas Indonesia. She holds a Bachelor’s degree in Biology and brings five years of professional experience in clinical laboratory settings prior to her current graduate studies. Currently in her second semester, her research focus lies at the intersection of metagenomics and skin microbiome dynamics. As an early researcher in the bioinformatics field, she is dedicated to learning more about computational biology analytical frameworks. She is also passionate about the practical applications of bioinformatics in health sciences and is actively expanding her expertise in genomic data analysis. This conference marks her debut as a scientific speaker, representing her formal entry into the international community.
A Reproducible Bioinformatics Workflow for Biomarker Discovery in Therapy-Resistant Breast Cancer Using Public Transcriptomic Datasets
- Presenter: Sepiah Dwi Cahyani
- Authors: Sepiah Cahyani (Universitas Indonesia)*
Abstract: Breast cancer resistance to endocrine and targeted therapies remains a major challenge in improving patient outcomes, particularly in hormone receptor-positive (HR+) subtypes. Public transcriptomic repositories provide valuable opportunities to investigate molecular mechanisms underlying therapy resistance through reproducible computational approaches. In this study, we developed a reproducible bioinformatics workflow to identify candidate biomarkers associated with therapy-resistant breast cancer using publicly available transcriptomic datasets, including GSE144378, GSE81620, and GSE222367 from the Gene Expression Omnibus. The workflow integrated differential gene expression analysis, cross-dataset comparison, functional enrichment analysis, and protein–protein interaction network construction using open-source bioinformatics tools and publicly accessible datasets. By combining transcriptomic profiles from tamoxifen-resistant and CDK4/6 inhibitor-resistant breast cancer models, we identified ten candidate genes consistently associated with resistant phenotypes across datasets. Functional analysis indicated that these genes are involved in pathways related to ECM-receptor interaction, focal adhesion, PI3K signaling pathway, and pathway in cancer. Beyond the biological findings, this study emphasizes the importance of reproducible computational workflows and FAIR principles in translational cancer research. The proposed pipeline demonstrates how publicly available transcriptomic datasets and interoperable open-source bioinformatics tools can be integrated to support scalable and reproducible biomarker discovery. This approach improves research accessibility across different resource settings and promotes the reuse of computational methods in the analysis of therapy-resistant breast cancer data.
Presenter Bio: Sepiah Dwi Cahyani is a Master’s student in Biomedical Science with a background in Biomedical Engineering. She has research experience in biomaterials and computational biology, with a growing interest in bioinformatics and reproducible research workflows. Her work includes transcriptomic data analysis for biomarker discovery in therapy-resistant breast cancer using publicly available datasets and open-source bioinformatics tools. She is particularly interested in integrating reproducible computational pipelines, FAIR data principles, and open science practices into biomedical research. Her current focus is on developing scalable and accessible bioinformatics workflows that support collaborative and cross-disciplinary research. She is passionate about bridging biomedical science and research software practices to improve transparency, reproducibility, and accessibility in scientific analysis.
Computational Transcriptomics Analysis for Biomarker Discovery in Triple Negative Breast Cancer
- Presenter: May Eva Francisca
- Authors: May Francisca (University of Indonesia)*
Abstract: Breast cancer (BC) is the second leading cause of cancer-related death among women worldwide, with Triple Negative Breast Cancer (TNBC) being the most aggressive subtype due to the absence of estrogen receptor (ER), progesterone receptor (PR), and HER2 expression. Despite its immunogenic characteristics, TNBC exhibits a high recurrence rate and poor prognosis, highlighting the need for more specific biomarker exploration. Identification of molecular biomarkers is essential for improving TNBC diagnosis, prognosis, and targeted therapy. This study employed a computational transcriptomics approach using publicly available gene expression datasets from the Gene Expression Omnibus (GEO). Three selected datasets were classified into TNBC and non-TNBC groups for Differentially Expressed Gene (DEG) analysis using GEO2R. DEGs were filtered based on |log2FC| ≥ 1. Protein–protein interaction (PPI) networks were constructed and analyzed using Cytoscape. Functional enrichment and KEGG pathway analyses were performed using ShinyGO to identify TNBC-related pathways and associated genes. Gene Ontology (GO) analysis was also conducted to evaluate biological processes, cellular components, and molecular functions associated with the identified genes. The analysis identified the top 10 potential biomarker genes associated with TNBC-related pathways, including pathways in cancer and breast cancer signaling pathways. These findings demonstrate that open-source bioinformatics software and reproducible computational workflows can support TNBC biomarker discovery and may contribute to future diagnostic, prognostic, and therapeutic development.
Presenter Bio: A Biomedical Engineering graduate currently pursuing a Master’s degree in Biomedical Science with a specialization in bioinformatics at Universitas Indonesia. My academic interests focus on bioinformatics, computational biology, biomolecular science, and translational biomedical research, particularly in integrating biological data analysis with healthcare applications. Through research and academic projects, I have developed experience in bioinformatics analysis, scientific data interpretation, and interdisciplinary problem-solving. I am passionate about leveraging computational approaches and collaborative research to support open, reproducible, and impactful scientific innovation. As an early-career researcher, I am eager to contribute to and learn from the global research software and bioinformatics community.
Scaling Reproducible Bioinformatics Workflows with R Markdown & Slurm
- Presenter: Felicia Ong Sing Yi
- Authors: Felicia Ong Sing Yi (Walter and Eliza Hall Institute of Medical Research)*
Abstract: Bioinformatics analyses often span multiple samples and large datasets, requiring scalable pipelines that are efficient, consistent, and reproducible. This talk explores workflow design principles for building standardized analytical pipelines using parameter flexibility, enabling a single template to be reused across different datasets while adapting to specific experimental conditions and analysis needs. In addition, RMarkdown caching reduces redundant computation by storing intermediate results, helping speed up reruns and improve efficiency in large-scale analyses where repeated execution is common. The Slurm workload manager further supports scalability by enabling the submission and management of parallel batch jobs, making it particularly effective for high-performance computing environments and multi-sample studies where tasks can be distributed across compute resources. Together, these approaches improve computational efficiency, reduce manual effort, and promote reproducibility. They also make analyses easier to adapt across different projects and datasets in bioinformatics research, ultimately making them easier to maintain in the long term.
Presenter Bio: Felicia Ong Sing Yi is a Research Assistant in the Chen Lab within the ACRF Cancer Biology and Stem Cells division, and the Bioinformatics and Computational Biology division at WEHI. She holds a Bachelor of Science with Honours (Distinction) in Computational Biology from the National University of Singapore. Felicia’s work centres on single-nuclei RNA and ATAC multi-omic data, where she develops analytical pipelines to interrogate gene regulation and chromatin remodelling pathways in cancer development. In this talk, she will share approaches for using RMarkdown caching and the Slurm workload manager to accelerate large-scale analyses, along with strategies for building standardized, reproducible workflows.
Reproducibility Debt: A Hidden Sustainability Challenge in Research Software
- Presenter: Zara Hassan
- Authors: Zara Hassan (Australian National University)*
Abstract: Reproducibility is fundamental to trustworthy and sustainable research software, yet many scientific software projects gradually become difficult to reproduce, maintain, and reuse over time. This presentation introduces Reproducibility Debt (RpD) as a new concept for understanding the long-term accumulation of reproducibility-related issues in research software ecosystems. Inspired by the idea of technical debt in software engineering, Reproducibility Debt describes how unresolved technical, organisational, documentation, workflow, infrastructure, and human-related issues accumulate over time and negatively impact the reproducibility and sustainability of scientific software. The presentation will demonstrate how seemingly small decisions, shortcuts, missing documentation, environment inconsistencies, dependency issues, poor collaboration practices, and lack of sustainable workflows can collectively create significant long-term reproducibility challenges. The talk will introduce a holistic taxonomy of Reproducibility Debt developed through empirical research on scientific software practices. Rather than viewing reproducibility purely as a technical issue, the taxonomy highlights the interconnected technical and non-technical dimensions that contribute to the accumulation of RpD across research software projects. The presentation will also emphasise the importance of Reproducibility Debt management and the need for sustainability-focused practices within research software communities. It will discuss how early identification, structured management approaches, and reproducibility-aware workflows can help reduce long-term maintenance burden, improve software sustainability, and support more reliable and reusable scientific research. This talk aims to provide research software practitioners, RSEs, and scientific communities with a new conceptual lens for understanding reproducibility challenges and supporting sustainable research software ecosystems.
Presenter Bio: Zara Hassan is a PhD researcher in Software Engineering at the Australian National University (ANU), specialising in reproducibility and sustainability challenges in scientific software. Her research introduced the concept of Reproducibility Debt (RpD), investigating how technical and non-technical factors accumulate over time to impact reproducibility in research software ecosystems. Her work focuses on developing holistic models and practical approaches for managing reproducibility challenges and improving the sustainability of scientific software.
Meta-valuation: A Self-Referential Mechanism for Valuing Diverse Research Contributions
- Presenter: Dr Cooper Smout
- Authors: Cooper Smout (Open Heart + Mind)*
Abstract: Research communities face a structural challenge: existing recognition systems reward outputs over process, fail to recognise diverse contribution types — including research software — and create misaligned incentives that undermine the very collaboration they depend on. Meta-valuation is a self-referential mechanism designed to address this, placing the full scope of contributions — including evaluations themselves — on a common scale, so that research software and other undervalued contributions can be given fair recognition and reward. In this talk, I’ll introduce the mechanism and present findings from recent experiments, including a live pilot at RSAA25 — where participants valued diverse conference contributions in real-time — and a more recent deployment at OHM Gathering 2025, where the mechanism coordinated and recognised contributions across a three-day in-person event. Collectively, these experiments demonstrate the mechanism’s capacity to value diverse contribution types — labour, capital, and ideas — and break down barriers between regions, institutions, disciplines, and communities, while remaining accessible to the diverse stakeholders who use it. I’ll also preview a next-generation valuation model inspired by these experiments, building on observed data patterns to improve value prediction while reducing reviewer burden — a critical barrier to broader adoption. Next, I’ll present theory on how the mechanism can address collective action problems in academia, helping communities transition to a more equitable, sustainable, and community-led research ecosystem. Communities who adopt the same, interoperable value dimensions can both coordinate activity and recognise the same contribution across contexts, creating a structural incentive for contributors to make their outputs openly available: sharing becomes the locally rational choice because it enables more recognition across contexts. Furthermore, the mechanism can recognise its own implementation and development, seeding a virtuous cycle of continuous improvement in which the valuation process becomes progressively more efficient. Together, these properties suggest that Open Science principles — such as openness, reliability, and replicability — may emerge naturally from the mechanism architecture alone, without need for external coercion. Rather, Open Science becomes the rational choice: as the most valuable form of contribution, it gets recognised and rewarded accordingly, aligning individual incentives with the common good.
Presenter Bio: Dr Cooper Smout is a designer, neuroscientist, and open science entrepreneur working at the intersection of collective intelligence, participatory governance, and cultural change. With a background in architecture and a PhD in the neuroscience of consciousness, he left academia to found Free Our Knowledge, a collective action platform for open research, and Open Heart & Mind (OHM), a nonprofit organisation developing open-source tools to empower communities. His current work focuses on meta-valuation, a participatory evaluation framework that generates transparent, interoperable value metrics to support fair recognition and reward, decentralised coordination, and collective governance across research and commons-oriented ecosystems.
Meta-valuation Across Borders: A Cross-Community Experiment in Valuing Diverse Research Contributions
- Presenter: Dr Cooper Smout
- Authors: Cooper Smout (Open Heart + Mind)*
Abstract: In this workshop we’ll explore meta-valuation, a novel mechanism for valuing diverse research contributions, and provide a proof-of-concept for its capacity to foster coordination and incentivise cross-border collaboration through interoperable value metrics. Its key innovation is a self-referential evaluation process, in which any contribution — including evaluations themselves — can be valued in relative terms, placing diverse research contributions on a common scale and aligning incentives toward the common good. At RSAA25, we conducted a live pilot test that produced relative valuations for diverse conference contributions in real-time, spanning presentations, financial sponsorship, facilitation, and behind-the-scenes efforts. Since then, a full deployment at OHM Gathering 2025 demonstrated the mechanism’s capacity to coordinate and recognise contributions across a multi-day community event. Together, these experiments demonstrate in principle that the same mechanism can value all forms of contribution — labour, capital, and ideas — through interoperable metrics tied to collective values. We now aim to take the next step: demonstrating interoperability in practice across a pair of workshops at RSAA26 and RSAfrica26. Each workshop will begin with a brief introduction to the mechanism and key findings. Participants will then nominate and approve contributions for review via a web-app interface, before collectively valuing contributions using a simple pairwise choice between randomly assigned contributions. The app will aggregate these votes into a relative value for each contribution and a total for each contributor. Crucially, both workshops will include local contributions — such as presentations — and global contributions — such as the common website infrastructure — enabling comparison and exploration of the mechanism’s interoperability across contexts. This will demonstrate a key property not yet evidenced: that recognition metrics can accumulate across communities who adopt the same evaluation dimensions, incentivising contributors to make their outputs openly available. We’ll close by celebrating the key contributors our data suggests made each conference happen, and reflecting on the process and results. Following both workshops, we will compare datasets, analyse findings, and aim to publish as a proof of concept for cross-community interoperability, with credit given to all contributors through the same mechanism explored within.
Presenter Bio: Dr. Cooper Smout is a designer, neuroscientist, and open science entrepreneur working at the intersection of collective intelligence, participatory governance, and cultural change. With a background in architecture and a PhD in the neuroscience of consciousness, he left academia to found Free Our Knowledge, a collective action platform for open research, and Open Heart & Mind (OHM), a nonprofit organisation developing open-source tools to empower communities. His current work focuses on meta-valuation, a participatory evaluation framework that generates transparent, interoperable value metrics to support fair recognition and reward, decentralised coordination, and collective governance across research and commons-oriented ecosystems.
Identification of Key Hub Genes in Hepatocellular Carcinoma via Integrated Bioinformatics Analysis and Stringent PPI Network Filtering
- Presenter: Tazkia Salma Hanifa
- Authors: Tazkia Hanifa (Universitas Indonesia)*
Abstract: Hepatocellular Carcinoma (HCC) remains one of the most aggressive malignancies with complex molecular mechanisms. This study aims to identify key hub genes involved in HCC progression through an integrated bioinformatics approach. The datasets were retrieved from the Gene Expression Omnibus (GEO) database and Differentially expressed genes (DEGs) were analyzed with GEO2R. The Protein-Protein Interaction (PPI) networks were constructed using the STRING and Cytoscape. Topological analysis was performed using the CytoHubba to identify the top 10 hub genes based on their connectivity scores. Functional enrichment analyses (GO and KEGG) were performed to elucidate the biological roles of these proteins. Three independent GEO datasets identified a total of 1,778 DEGs. After mapping these genes to the STRING database, only 1,736 DEGs were successfully recognized. To ensure high-confidence interactions, the network was filtered using a stringent interaction score threshold of ≥ 0.999, resulting in a core subset of 300 genes. Gene Ontology (GO) enrichment analysis of these genes revealed significant involvement in the regulation of apoptotic processes (Biological Process), association with the cyclin D2-CDK4 and Bcl-2 family protein complexes (Cellular Component), and high activity in BH3 domain binding and DNA-binding transcription activator activity (Molecular Function). Subsequent pathway analysis via ShinyGO prioritized the “Pathways in Cancer” category, from which 26 genes were curated for further refinement. Finally, topological analysis using CytoHubba identified 10 key hub genes: CDK4, TP53, FOS, HSP90AA1, AR, BCL2L1, FOXO1, BCL2, ESR1, and CDKN2A, which serve as the primary drivers in HCC interaction network. The findings highlight a specific cluster of hub genes that drive hepatocarcinogenesis. Those 10 genes may serve as robust candidate biomarkers for the diagnosis and prognosis of Hepatocellular Carcinoma.
Presenter Bio: A biomedical researcher and bioinformatician working at the intersection of multi-omics data—specifically metabolomics and transcriptomics—and scientific computing for biomarker discovery and drug target identification. She heavily leverages open-source software ecosystems to translate complex biological data into actionable medical insights. Aligned with the spirit of Research Software Without Borders, She firmly believes about how technology and open-source tools can accelerate scientific innovation and global health solutions in a collaborative global scientific community.
Reproducible command-line workflows with asciinema-scripted
- Presenter: Robert Moss
- Authors: Robert Moss (The University of Melbourne)*
Abstract: A picture is worth a thousand words. But screenshots and video recordings of text-based workflows may not always be particularly helpful, because their pictorial format makes the text inaccessible to selecting, copying, and pasting, and may even render the text blurry and illegible. One alternative is asciinema (https://asciinema.org/), which records terminal sessions and plays them back in plain text, allowing the viewer to pause and copy/paste any of the content. While this can be beneficial, particularly for educational purposes, it also captures every typographic error and unintended pause during the recording. For this reason, I have developed a simple tool (asciinema-scripted) for scripting terminal sessions that provides control over input speed, and supports chapter markers and inline comments/subtitles. I have used this tool to create several recordings for our online training materials (Git is my lab book) and have received unprompted feedback from several users who found this to be a surprising and useful feature. In this presentation I will give a brief overview of this tool and show a recording that demonstrates its features.
Presenter Bio: Robert Moss is a Senior Research Fellow in the Centre for Epidemiology and Biostatistics at the Melbourne School of Population and Global Health, and a member of Australia’s National Respiratory Infections Surveillance Committee. His research focuses on predicting and mitigating the burden of infectious diseases, through the use of mathematical/computational modelling and Bayesian inference. He is a proponent of reproducible research, has developed online training materials for researchers (https://git-is-my-lab-book.net/), and has produced and maintains a number of open source software packages. Contact him via email ([email protected]) or mastodon (@[email protected]).
Breaking Bias in Requirements Engineering: Artificial Intelligence for Inclusivity & Sustainability
- Presenter: Dr Hafsa Shareef Dar
- Authors: Hafsa Dar (University of Gujrat)*
Abstract: Requirements Engineering (RE) is the foundational phase of software development, where biases embedded in natural language (NL) requirements can propagate throughout the software lifecycle, potentially resulting in systems that exclude or disadvantage certain user groups. Despite increasing awareness of Diversity and Inclusion (D&I) and the United Nations Sustainable Development Goals (SDGs), the integration of social sustainability—particularly SDG-5 (Gender Equality)—into RE remains largely theoretical, with limited practical tools, methods, and AI-based support for bias detection during requirements elicitation. Key gaps in current RE practices include the lack of systematic mechanisms to detect gendered, stereotypical, or exclusionary language; insufficient analysis of stakeholder diversity and representation; absence of SDG-aligned evaluation metrics; and reliance on manual, ad-hoc bias mitigation approaches. Traditional RE reviews are often subjective, time-consuming, and ineffective in capturing subtle forms of bias related to gender, culture, or accessibility. Moreover, sustainability efforts in software engineering have primarily focused on environmental and economic aspects, while neglecting the social dimension. To address these limitations, this research proposes an AI-powered NLP framework for inclusive requirements engineering. The framework applies natural language processing and machine learning techniques to automatically detect gendered language, stereotypes, and exclusionary expressions in requirements documents. It also identifies gaps in stakeholder representation and evaluates requirements alignment with SDG-5 targets and broader social sustainability goals. The framework further provides actionable feedback for bias mitigation and inclusive rewriting of requirements. Embedded mitigation strategies include gender sensitization support, generation of inclusive personas, development of inclusive conceptual models, and iterative feedback-driven refinement of requirements specifications. Unlike prior exploratory work, this approach emphasizes practical implementation supported by tools and measurable evaluation criteria. Overall, the proposed solution promotes Inclusive-by-Design Requirements Engineering by enabling fair representation, transparent bias detection, and improved equity in software outcomes. It supports the development of socially responsible and sustainable software systems that enhance access, participation, and empowerment across diverse user groups. By integrating D&I considerations early in the development lifecycle, the framework aligns software engineering practice with the UN 2030 Agenda. This research contributes by bridging AI, Requirements Engineering, and sustainability studies, offering practical methods, processes, and metrics that have been largely missing.
Presenter Bio: Dr. Hafsa Dar is a PhD in Software Engineering with over 15 years of experience in academia and industry. She earned her BS, MS, and PhD from International Islamic University Islamabad, Pakistan. Her research focuses on Requirements Engineering, Diversity and Inclusion (D&I), and AI/NLP-based bias detection for sustainable software systems. She is particularly interested in integrating social sustainability, especially SDG-5 (Gender Equality), into Requirements Engineering through inclusive-by-design approaches. She has published in peer-reviewed journals and presented at national and international conferences, and serves as a reviewer and program committee member. She has contributed to collaborative research with Middlesex University and University of Hertfordshire on sustainability in Requirements Engineering. She is Co-Director of Women in Big Data (Pakistan), promoting skills development for women in data science and software engineering. She was selected for the IVLP 2019 (USA).
Life Registries and the Unborn Child: A Matter of Data Access and Protection?
- Presenter: Dennis Batangan
- Authors: Dennis Batangan (Ateneo de Manila University, Institute of Philippine Culture)*
Abstract: The first one thousand (1000) days of life is a very critical period in the human body’s growth and development when it undergoes a continuous development process in response to biological, genetic, and environmental stimuli. These growth and development processes, in the longer term. establishes the bases for a healthy and productive life. However, countries like the Philippines still have policy and programmatic gaps in the civil registration and health information system that can provide the inputs for more proactive and preventive health interventions for the unborn child. A key element of the said information system is the baseline health information from birth and a neonatal clinical history to connect the information of the newborn to the data of mother’s pregnancy, childbirth, infancy, adolescence, and adulthood. In many countries, including the Philippines, there are no explicit laws pertaining to whether the unborn child can be included in life registries and information systems specially since ‘it is birth that determines personality’ under the Civil Code of the Philippines (Republic Act Number 386) ,Although information about the unborn child is considered health-related data, which are deemed sensitive personal information under the Philippine Data Protection Act of 2012(Republic Act 10173) , the protection afforded by the law is not explicit and there remains outstanding ethical and legal issues which remain unresolved. Reviewing relevant international covenants like the Global Digital Compact, World Health Organization (WHO) guidelines, Philippine jurisprudence and precedents from Spain and Japan, the presentation will interrogate the issues around the rights of the unborn child and the authority of the state to collect and process the data of an unborn child.
Presenter Bio: Dr. Batangan is the 2025 Lourdes Campos Awardee for Public Health given by the Philippine Association for the Advancement of Science and Technology (PhiAAST). He is a distinguished faculty member and research scientist at Ateneo de Manila University. He holds positions within the School of Social Sciences, School of Government, and Institute of Philippine Culture. With degrees in psychology and medicine from the University of the Philippines, he furthered his studies with a post-graduate course in Community Health for Developing Countries at the University of Heidelberg, Germany. Dr. Batangan’s career is marked by transformative contributions to community and public health. He co-founded the Community Medicine Development Foundation, Inc (COMMED), pioneering social innovations. His dedication earned him several community health service awards in the Philippines.
From Archive to Analysis: Community-Led Open Software for HASS, Indigenous and GLAM Digital Collections
- Presenter: Natalia Orel
- Authors: Natalia Orel (University of Melbourne)*
Abstract: Australia’s galleries, libraries, archives, and museums hold remarkable records of our history, culture, and language, with much of it available online. However, having access to digitised collections is only half the story. Researchers still lack the practical tools and training to work with these materials computationally, and the gap between “digitised” and “usable for research” remains wide. This presentation shares work-in-progress from the Enhanced Analytics project: Melbourne Data Analytics Platform’s (MDAP’s) contribution to Phase 2 of the ARDC Community Data Lab. We’re building open-source software to address specific, researcher-identified gaps in how HASS&I and GLAM communities can access, analyse and get value from digital collections.
Presenter Bio: Natalia is a Research Software Engineer at the Melbourne Data Analytics Platform (MDAP) at the University of Melbourne. Her current work focuses on open-source Python tools for HASS and GLAM research data, as part of the ARDC Community Data Lab Phase 2 Enhanced Analytics project. Before moving into research software, she worked across biotechnology and intellectual property analysis.
Transitioning to Minimally Invasive Diagnostics for Atopic Dermatitis: Validating a Keratin-Based Genetic Signature in Skin Shave Samples
- Presenter: Rachmat Puziyanto
- Authors: Rachmat Puziyanto (University of Indonesia)*
Abstract: Background: Although skin biopsies are the gold standard for molecular studies in Atopic Dermatitis (AD), their invasive nature often leads to patient discomfort and permanent scarring. Consequently, there is a pressing need for diagnostic methods that are less traumatic yet maintain high molecular precision. In this study, we aimed to bridge this gap by identifying a stable genetic signature from biopsy data and testing its reliability in superficial skin shave samples. Methods: We implemented an integrative bioinformatics workflow, starting with a meta-analysis of two independent biopsy datasets from Gene Expression Omnibus Database (GSE5667 and GSE32924) to isolate consistent Differentially Expressed Genes (DEGs). These genes were then filtered through functional enrichment and Protein-Protein Interaction (PPI) mapping, using a stringent confidence threshold (>0.900) to pinpoint the most critical hub genes. Finally, we validated the performance of this signature on a third, independent dataset (GSE60709) collected via skin shaving, using SVM, Random Forest, and Naive Bayes classifiers. Results: Our analysis identified a core signature consisting of three epidermal stress-response genes KRT16, KRT6A, and KRT6B acting as key biomarkers. Validation results were highly encouraging; this biomarker trio effectively distinguished between lesional and non-lesional skin in shave samples, with the SVM model achieving an outstanding AUC of 0.929. Both Random Forest and Naive Bayes models also showed high accuracy with an AUC of 0.917. Conclusion: These results demonstrate that the KRT16, KRT6A, and KRT6B biomarker signature remains remarkably stable even when moving from deep-tissue biopsies to superficial skin shaves. By maintaining such high diagnostic accuracy, our approach offers a viable path toward replacing invasive biopsies with patient-friendly sampling, potentially transforming how we monitor and diagnose Atopic Dermatitis in routine clinical practice.
Presenter Bio: My name is Rachmat Puziyanto. I am a medical doctor currently pursuing a Master’s in Biomedical Sciences at the University of Indonesia, Faculty of Medicine, where I specialize in Bioinformatics. My work sits at the intersection of clinical practice and computational biology, where I focus on making advanced molecular diagnostics more accessible and patient-friendly. I am particularly passionate about transitioning from invasive procedures to minimally invasive methods. My current research involves using bioinformatics and machine learning to validate genetic signatures for Atopic Dermatitis via skin shave samples aiming to replace painful biopsies with a much simpler, scar-free alternative. As a clinician navigating the world of research software, I see myself as a bridge between the bedside and the data lab.
Birds of a Feather: RSE-AUNZ Update, Community Survey & Ask Me Anything
- Presenters: Rowland Mosbergen, Matthew Laurenson
- Authors: Rowland Mosbergen (WEHI)*; Matthew Laurenson (Bioeconomy Science Institute)
Abstract: Join us for an interactive session where we’ll share some accomplishments of the Research Software Engineer Community of Australia and New Zealand (RSE-AUNZ), talk about community needs, host an Ask Me Anything session, and collect and develop ideas for our 2026–2030 strategy. This session welcomes experienced research software engineers, early-career professionals, research software team leaders, and institutional supporters of RSE practice. We will present an overview of RSE-AUNZ’s recent initiatives, including Zulip, our LinkedIn page, and our membership database, and their impact. We’ll follow that with a live community survey designed to gather insights into your priorities, organisational challenges, and strategic vision for research software engineering across the Australia / New Zealand region. An Ask Me Anything (AMA) session with senior RSE leaders, that will spark discussion about where we should focus our efforts over the next four years. This discussion will allow attendees to contribute to RSE-AUNZ’s strategic direction while engaging with others navigating similar professional challenges and opportunities. In the second session, we will break into groups to prioritise the ideas resulting from the AMA, and the next steps on how to proceed with actions following the conference.
Presenter Bio: Rowland is a strategic leader in digital research infrastructure with 25+ years’ experience translating HPC, AI, data and cloud technologies into scalable research and industry services used by global communities. Known for building partnerships across researchers, national infrastructure providers and technology teams to deliver high-impact digital capability, Rowland provides strategic direction, inclusive leadership and scalable programs that accelerate research adoption and demonstrate measurable impact. Rowland co-founded RSAA, RSAfrica, and RS Latin America, and established Equersa, the umbrella organisation supporting all three regional conferences. He also serves on the steering committee of the RSE Australia New Zealand Association.
Building Diverse and Inclusive Technical Teams: From Recruitment to Retention
- Presenters: Rowland Mosbergen, Marion Weinzierl
- Authors: Rowland Mosbergen (WEHI)*; Marion Weinzierl (Institute of Computing for Climate Science (ICCS))
Abstract: Technical teams - whether in research or elsewhere - miss out on talent and creativity because of a lack of diversity and inclusivity. And it is very hard to change that. Diversity Equity and Inclusion (DEI) checklists can help get a foot in the door, but improving team diversity and accessibility is a marathon, not a sprint. This workshop helps participants build more accessible, inclusive, and diverse teams. You’ll learn practical strategies for recruitment, get hands-on guidance to improve your process, and discover how to maintain this approach for employee retention. The target audience are people recruiting for technical teams, specifically Research Technical Professionals such as Research Software Engineers, Research Platform Engineers, Data Managers, Digital Lab Technicians, etc.
Presenter Bio: Rowland: Rowland is a strategic leader in digital research infrastructure with 25+ years’ experience translating HPC, AI, data and cloud technologies into scalable research and industry services used by global communities. Known for building partnerships across researchers, national infrastructure providers and technology teams to deliver high-impact digital capability, Rowland provides strategic direction, inclusive leadership and scalable programs that accelerate research adoption and demonstrate measurable impact. Rowland co-founded RSAA, RSAfrica, and RS Latin America, and established Equersa, the umbrella organisation supporting all three regional conferences. He also serves on the steering committee of the RSE Australia New Zealand Association. Marion: Marion Weinzierl received the M.Sc. (German Diploma) degree in media informatics from Ulm University, Germany, in 2008, and the Ph.D. (Dr.rer.nat.) degree in scientific computing from the Technical University of Munich, Germany, in 2013. She was a Postdoctoral Research Associate in computational solar physics with Durham University, U.K., from 2014 to 2017, and a Computational Scientist and a Senior Computational Scientist with IBEX Innovations, Sedgefield, U.K., from 2017 to 2018 and from 2018 to 2019. From 2019 to 2022 and from 2022 to 2023, she was a Research Software Engineer in advanced research computing and a Senior Research Software Engineer with Durham University, U.K. From 2020 to 2023, she was also the Research Software Engineering Theme Leader of the N8 Centre of Excellence in Computationally Intensive Research (N8 CIR). Since 2023, she has been a Senior Research Software Engineer with the Institute of Computing for Climate Science, University of Cambridge, U.K.
Open Source as Public Infrastructure: Early Evidence from Africa
- Presenters: Tosan Okome, Chidera Frankie
- Authors: Tosan okome (CHAOSS Africa)*; Chidera Frankie (CHAOSS Africa)
Abstract: Open source software (OSS) is quietly powering key public systems across Africa. Yet governments rarely recognise it as part of their Digital Public Infrastructure (DPI). Tools like DHIS2, OpenMRS, Mojaloop, and Moodle support health data management, financial inclusion, and education delivery in dozens of African countries. But formal policy frameworks, sustained funding, and institutional recognition remain low or absent in most contexts. This presentation shares early findings from research being conducted by the CHAOSS Africa Research Focus Group. We are a small, community-led team studying the health and sustainability of open-source communities across the continent. Our research examines the adoption and implementation of open source technologies across African DPI deployments, assesses their documented impact (cost savings, efficiency gains, accessibility improvements), reviews global and continental policy frameworks, and identifies actionable recommendations for African governments. The research is explicitly work-in-progress. We will share what we have found, what remains unresolved, and how the research software community can contribute, whether through data, critique, or collaboration. This talk connects somewhat directly to the conference theme of Research Software Without Borders. The OSS tools we study (such as DHIS2) are research software in the broadest sense; built collaboratively, maintained by distributed communities, and designed for reuse across institutional and national borders. Their sustainability challenges mirror those facing research software generally: funding gaps after initial deployment, limited local maintainer capacity, and weak institutional recognition. Our work also models community-led research practice in the Global South — conducted by a small team of novice but motivated researchers, organised through CHAOSS Africa’s open community infrastructure, and aimed at producing outputs (a report and a policy playbook) that remain accessible and actionable beyond the academic context.
Presenter Bio: Chidera Frankie and Tosan Okome. Chidera Frankie is a software engineer with a strong curiosity for research, particularly in how technology shapes society and public systems. She is the project manager for the ongoing research by the CHAOSS Africa Research Focus Group on Open Source as digital public infrastructure. Tosan Okome is a registered nurse, midwife, and public health nurse based in Nigeria with interests in research, data analytics, and open science. She leads the Research Focus Group within CHAOSS Africa and serves as Director of Research for the Pre-Seeds Project, an introductory research programme supporting underrepresented individuals. Tosan is passionate about open, reproducible research and the role of data in strengthening research ecosystems across Africa.
HPC Resource Dashboard: Helping Researchers Do More With Less on Shared HPC Infrastructure
- Presenter: Samuel Green
- Authors: Samuel Green (UNSW)*
Abstract: High-performance computing allocations are a finite shared resource, yet most researchers have limited visibility into how efficiently their jobs use what they request. We present the HPC Resource Dashboard - an open-source, self-hosted web platform that gives researchers real-time and historical insight into the CPU, memory, GPU, and I/O usage of their compute jobs on NCI’s Gadi supercomputer. A lightweight Rust telemetry agent (1–2% CPU overhead) runs alongside each job and writes metrics to a shared filesystem. A Nirin cloud-hosted backend ingests these logs automatically and serves interactive dashboards showing per-job resource efficiency, time-series charts, per-core CPU breakdowns, and automatic recommendations for right-sizing future submissions. The system supports PBS and Slurm schedulers and handles 50–100 concurrent users on a single modest VM (8 vCPU, 16 GB RAM) through seek-based file tailing, query-level caching, and container memory limits. In this demo, we will walk through the full workflow: instrumenting a Gadi job with one line of shell script, watching telemetry stream into the dashboard in real time, interpreting the efficiency metrics, and using the automatic resource recommendations to reduce waste on subsequent runs. We will also discuss the architecture decisions that allow the platform to operate within the constrained resources of a research cloud allocation , a practical example of doing more with limited infrastructure. The platform is live at hpc-dash.au, fully open source, and designed to be deployable at any site running PBS or Slurm with a shared filesystem.
Presenter Bio: Samuel Green is a Research Software Engineer at UNSW Sydney and the ARC Centre of Excellence for 21st Century Weather, working in the fields of climate science, high-performance computing, and scientific data engineering. His current work spans large-scale atmospheric reanalysis pipelines, machine learning for environmental applications, and scientific visualisation - with a particular focus on making HPC workflows more efficient and reproducible on systems like NCI’s Gadi and Pawsey’s Setonix. He developed the HPC Resource Dashboard to address a practical problem he observed across research groups at NCI: a lack of visibility into how compute resources are actually used. Sam has a background in astrophysics where his previous research included predicting thermal radiation from stellar wind bubbles using multi-dimensional simulations. Sam holds a Bachelor degree in Physics & Astrophysics, a Masters degree in Space Science, and a PhD in Computational Astrophysics.
Dual-Purpose Research Software
- Presenter: Attila Egri-Nagy
- Authors: Attila Egri-Nagy (Akita International University)*
Abstract: For the ancient game of Go, it is a common wisdom that you have to make moves that serve more than one goal. In this talk, we apply this golden rule to developing research software. We will tell the story of how an educational opportunity was spotted during the development of a mathematical (algebraic automata and category theory) software package. The original goal was to have a programming language that has a syntax so simple, that it can be a direct input to abstract algebraic decomposition algorithms. The design of the language was guided by actual programming in it. Writing those test program turned out to be entertaining. The minimalist syntax impose the type of constraints that induce creativity much like the microcomputers of the 80’s are still used for recreational coding. Another benefit extreme parsimony of the core language is that so little needs to be explained to get started. These are the hallmarks of a language suitable for complete beginners. The “coding for all” movements are often motivated by the possible economic benefits. These may be reduced by the progress in automated software engineering (e.g., agentic coding). We argue that the experiential side of programming is becoming more relevant. Perhaps, with the long-term goal of promoting and retaining an understanding of computing devices, as part of human culture. Therefore, we integrated interactive sessions of learning programming into a liberal arts program. The talk will briefly introduce the developed concatenative functional programming language and will report on the teaching sessions.
Presenter Bio: Attila is mathematician turned software engineer working in algebraic and category theoretical investigations of the nature of computing. He still handcrafts code, mainly in Clojure. He also teaches a wide range of courses in mathematics and computer science with a passion for a culture where thinking and problem solving is a fun, social and meaningful process.
Australian BioCommons: infrastructure access and execution for complex research workflows
- Presenters: Johan Gustafsson, Steven Manos
- Authors: Johan Gustafsson (University of Melbourne)*; Steven Manos (University of Melbourne)
Abstract: Australian BioCommons is the national digital research infrastructure for molecular life science researchers across academia, government, and industry. BioCommons responds to priority challenges in bioinformatics that are identified through consultation with Australian molecular life science communities (e.g. microbiome, proteomics), research consortia, and national projects. These responses are built on the real-world experience of researchers, and directly address the challenges faced by adopting, aligning to, or deploying national platforms and services.
In this presentation, we will provide an update on BioCommons tailored to the research software engineering (RSE) community. It will cover the challenges we are addressing to improve access and execution of computational software and research workflows, and how our efforts to integrate elements of the Australian and international infrastructure ecosystem map to both the software life cycle and the requirements of RSEs.
This talk aligns with multiple topics for the conference, including:
- Community leadership, cross-institution and cross-regional collaboration
- Real-world experiences and lessons learned
- Open collaboration and reuse of shared platforms, that use a tailored combination of cloud and HPC
- Use of open, sovereign infrastructure, where possible, combined with existing international platforms and standards
- Development of reproducible workflows that embed good research practice, including FAIR software in practice
- Teaching and learning research software skills
- Supporting emerging technologies, including AI-enabled research software and workflows
Presenter Bio: Dr Johan Gustafsson is a Research Community Engagement Lead at Australian BioCommons, focused on computational methods and workflows across proteomics, metabolomics, and structural biology. He forges community connections both nationally and internationally, works closely with communities to develop roadmaps that present a vision for shared national infrastructure, and coordinates on-going consultations that help to roll out their shared vision. https://orcid.org/0000-0002-2977-5032.
Making space: Celebrating people and diversity in Research Software
- Presenter: Jana Makar
- Authors: Jana Makar (Women in HPC Australasia Chapter)*
Abstract: Visibility matters. As the Australasia Chapter of Women in High Performance Computing (WHPC+ AusNZ) we aim to raise the visibility and celebrate the successes of underrepresented groups in fields of computational research and research software. Fitting with RSAA’s 2026 theme “Research Software Without Borders,” this session will share different career pathways in research software to foster dialogue around lived experiences and entry points into the field. We will create space for an open conversation on how we can collectively break down barriers to entry and build a more inclusive research software community. This session will be split into two parts: 1) A ‘showcase’ of lightning talks by WHPC+ Chapter community members who are working on research software projects and the career paths they’ve had to get to the roles they’re in today. Our guest speakers for this part are: Anjusree Karnavar: PhD student, Griffith’s University, Gayatri Aniruddha: Associate lecturer, University of Western Australia. 2) An open discussion around actionable steps our communities can take to create a more diverse and supportive ecosystem for people working in research software. Key insights will be captured and shared in a post-session summary to help carry these conversations beyond the session.
Presenter Bio: Facilitator: Women in HPC Australasia. Guest speakers: Anjusree Karnavar: PhD student, Griffith’s University, Gayatri Aniruddha: Associate lecturer, University of Western Australia. Anjusree Karnavar: I am a Doctoral Researcher at Griffith University, Queensland, Australia, working in collaboration with the Commonwealth Scientific and Industrial Research Organization (CSIRO). My research focuses on Computer Vision, Image Restoration, and Deep Learning, developing advanced AI techniques, using diffusion models and Vision Transformers, to restore and enhance degraded images for more reliable vision-based systems. Before pursuing my PhD, I worked for six years as an Assistant Professor in Computer Science and Engineering at APJ Abdul Kalam Technological University, India. My interests lie in advancing AI-driven methods for visual computing and intelligent systems. Gayatri Aniruddha: (bio coming soon!)
Bridging the Gap: BoF for Standardising User and Dataset Mapping Across Research Software Ecosystems
- Presenters: Rowland Mosbergen, Anurag Katariya, Melroy Almeida, Sarah Thomas
- Authors: Mosbergen, Rowland; Katariya, Anurag; Almeida , Melroy; Thomas, Sarah*
Abstract: Facilitated by the Australian Access Federation, this Birds of a Feather session will explore a critical challenge in research software ecosystems: managing user access and dataset permissions across multiple specialised applications. The Australian Access Federation (AAF) is committed to transforming Australia’s research, teaching, and learning communities - they do this by delivering innovative technology solutions and policy, that support expert, secure access to digital resources and infrastructure, across the entire teaching and research ecosystem. Importantly, AAF is a core part of the Australian National Digital Research Infrastructure, serving as the Trust and Identity capability that enables trusted access to Australia’s leading research infrastructures. Australian researchers rely on a diverse range of commercial and open-source software and infrastructure. Many key research applications, data repositories, and infrastructures have limited or no ability to leverage the foundational building blocks of system-wide T&I solutions for research. When organisations deploy diverse specialised research applications, like cBioPortal, REDCap, Vitessce, Stemformatics, and Degust, there is no standardised way to consistently represent researchers, and their data access rights across those systems. As a result, researchers face fragmented access experiences rather than seamless, single sign-on like workflows, while organisations struggle to apply consistent dataset-level access controls across applications. Through presentations from organisations that have tackled this challenge and facilitated group discussion by technical and policy experts from the AAF, we will document integration strategies, identify common patterns, and work towards developing open standards and identifying opportunities for shared libraries that the research software community can adopt.
Presenter Bio: Rowland Mosbergen is a strategic leader in digital research infrastructure with 25+ years’ experience translating HPC, AI, data and cloud technologies into scalable research and industry services used by global communities. Known for building partnerships across researchers, national infrastructure providers and technology teams to deliver high-impact digital capability, Rowland provides strategic direction, inclusive leadership and scalable programs that accelerate research adoption and demonstrate measurable impact. Anurag Katariya is a Technical Consultant at the Australian Access Federation, with experience in Identity and Access Management, federated identity and solutions architecture. He has a particular interest in federated access, authorisation and persistent identifiers and is passionate about improving the researcher experience through interoperable Trust and Identity solutions across Australia’s research ecosystem. Melroy Almeida is a Portfolio Manager – PIDs at the Australian Access Federation. He has a strong interest in persistent identifiers and recognises their value in enhancing research and promoting FAIR principles. He also collaborates with peers to advance the adoption of persistent identifiers and foster a connected community of practice. Sarah Thomas is a Portfolio Manager at the Australian Access Federation. She has been working on the delivery of Trust and Identity for national research infrastructure, partnering with the community on co-designing a unified approach to trusted and seamless access for Australia’s research sector. An enthusiastic and dedicated community engagement manager, with demonstrated expertise in research and innovation ecosystems. Committed to using her skills in relationship management, communication, program design, delivery and evaluation to deliver positive outcomes for the community.
BoF: Metadata Across Disciplines – A Collaborative Journey
- Presenters: Paul Wang, Emily Fitzgerald, Rowland Mosbergen, Scout Bell
- Authors: Mosbergen, Rowland*; Wang, Paul; Fitzgerald, Emily; Bell, Scout
Abstract: What metadata problems keep you up at night? Join us to find allies, discover existing tools you might not know about, and build a community around this messier-than-we-admit reality. Metadata management is a universal challenge, yet solutions remain siloed within disciplines. This Birds of a Feather session brings together humanities scholars, social scientists, clinicians, archivists, and data stewards to explore metadata lifecycle problems through a shared lens - not to prescribe solutions, but to discover common ground. We’ll sample some of the biggest challenges that a HASS researcher and a Life Science researcher face: they might be challenges from collection challenges through to data cleaning, provenance tracking, and structural decisions (flexibility versus rigor). This session features two concrete high-level case studies - one from humanities and social sciences, one from life sciences - showing how metadata problems manifest differently yet share surprising patterns. Participants will identify the three biggest bottlenecks in their own domains and map where they occur in the lifecycle. We want the community to share practical synthetic metadata that is based on real life (both clean and messy), to use as test criteria for evaluating solutions. And we will crowdsource this information into a living document: a high-level guide that attendees can use as a reference in the future.
Presenter Bio: Emily Fitzgerald is a historian and Research Data Specialist in the Melbourne Data Analytics Platform (MDAP) at the University of Melbourne. Her PhD research on the transnational connection between Australia and the United States during the development of Australian Federation fostered her interested in the digital humanities, where she has a particular affection for creating maps and data cleaning. She is keenly interested in what technological tools and skills can bring to humanities research and what humanities scholars can bring to technological development. Scout Bell (they/them) (also known as Emilia C. Bell) is a PhD candidate at Curtin University and Manager, Research and Digital Services at Murdoch University Library. Their doctoral research explores open knowledge diplomacy, focusing on Australian university libraries in open practices for climate scholarship. With experience in organisational leadership, evidence-based practice, research services, and disability and neurodiversity inclusion, Scout delivers collaborative and transformational outcomes. In addition to their research, they serve as Board Director for Ballet Without Borders, previously served on the Board of the Australian Library and Information Association (ALIA), and are Co-Founder and President of the Alliance for Neurodiversity and Disability in GLAMR Professions Australia Inc. (ANDPA).
From Ad Hoc Promotions to Visible Career Pathways for Research Software Engineers
- Presenters: Rowland Mosbergen, Julie Iskander
- Authors: Mosbergen, Rowland*; Iskander, Julie
Abstract: Research Software Engineers are often described as lacking visible career pathways, but the practical meaning of this problem is often under-specified – so what does this mean in practice? RSE roles encompass diverse specialisations—data engineers, HPC specialists, and web application developers—yet institutional career structures frequently fail to recognise these distinctions. Classification, approval, and promotion processes can be slow when role expectations are not clearly defined in advance. At the WEHI Medical Research Institute, we addressed this challenge proactively by designing a comprehensive, multi-level career framework tailored to RSE specialisations. We developed distinct position descriptions for Research Data Engineers, Research Compute Engineers (HPC), and Research Software Engineers, each with three clearly defined career levels: Junior Engineer, Engineer, and Senior Engineer - with an overarching Principal Engineer. This tiered framework makes career progression more transparent by clarifying expectations, responsibilities, and growth pathways across different technical specialisations. It also provides a stronger basis for recruitment, promotion discussions, capability planning, and long-term retention of skilled research technology staff. This presentation will share the rationale, design process, challenges, and lessons learnt from developing the framework at WEHI. It will also discuss how similar institutions can adapt the approach to their own size, structure, and maturity. By moving from reactive role design to proactive career planning, research organisations can strengthen their research software and computing capability while making technical careers more visible, credible, and sustainable.
Presenter Bio: Julie Iskander and Rowland Mosbergen (Bios provided above).
Building and sustaining research software engineering communities: The role of computational skills trainers
- Presenter: Elijah Atta Manu
- Authors: Atta Manu, Elijah*
Abstract: Research has been the backbone of various significant innovations, yielding enormous societal and financial benefits. Modern research has become increasingly reliant on large-scale and complex data acquisition and analyses. These often involve cloud storage and complex computational workflows. These requirements for modern research underscore the important role played by Research Software Engineers in the development and deployment of powerful software, whose effective adoption will potentially drive further innovation. To realise these benefits, it is important to ensure the sustainability of a community that not only develops research software, but also champions the effective adoption of these software by researchers. While product usage demonstrations may be routinely undertaken by research equipment and software developers, researchers still face several barriers that often result from limited computational skills such as command-line basics and debugging intuition. Effective adoption of new research software and the sustainability of the Software Engineering community both depend on educating researchers. This education must cover the foundational principles on which research software are developed, and the logical reasoning behind how they operate to deliver desired outcomes. This talk discusses the crucial role of computational skills trainers in sustaining the research software engineering community, drawing on lessons from The Carpentries, inc. and their instructor training programme.
Presenter Bio: I am a final-year PhD Candidate at the University of Otago, New Zealand, and an Assistant Lecturer of Physiology at the Kwame Nkrumah University of Science and Technology, Ghana. My research areas of expertise are Cell and Molecular Physiology and Cancer. I am also a certified Carpentries instructor, and volunteer as a Training Affiliate with Genomics Aotearoa (GA). Together with the training team at GA and their partners at Research and Education Advanced Network New Zealand (REANZ), I contribute to the upskilling of New Zealand researchers through various bioinformatics workshops.
From R to Python: Porting a Research Software Package Across Language Ecosystems
- Presenter: Jyoti Bhogal
- Authors: Bhogal, Jyoti*
Abstract: Research software is often developed within a single programming language ecosystem, which can limit its accessibility and adoption by broader research communities. This presentation shares the journey of taking an R package, ‘generatervis’, and developing a Python implementation to make its functionality available to a wider audience. ‘generatervis’ was originally created to simplify common data visualisation tasks and reduce repetitive coding in research workflows. As interest in the project grew, a natural question emerged: what would it take to support researchers who primarily work in Python? This talk will explore the process of translating a research software project across programming languages while preserving its core design principles and user experience. It will cover technical decisions, package architecture, testing strategies, documentation, dependency management, and publishing workflows. The presentation will also discuss the similarities and differences between the R and Python packaging ecosystems, including lessons learned while working with modern Python development tools. Beyond the technical aspects, this experience highlights a broader challenge in research software: making tools accessible across disciplinary and technical boundaries. By sharing successes, mistakes, and practical lessons from this cross-language development effort, the talk aims to help other research software developers who are considering expanding their tools beyond a single language ecosystem. Attendees will gain insights into software portability, sustainable package development, and practical approaches for building research software that can reach more diverse research communities. Project Resources: 1) ‘generatervis’ GitHub Repository: https://github.com/Clinical-Informatics-Collaborative/generatervis 2) Development Blog Post 1: https://jyoti-bhogal.github.io/about-me/blog/2025-04-17_generatervis_blog_1/ 3) Development Blog Post 2: https://jyoti-bhogal.github.io/about-me/blog/2025-04-29_generatervis_blog_2/
Presenter Bio: Jyoti Bhogal is a statistician, community builder, and advocate for Research Software Engineering (RSE). She is Co-Lead of the RSE Asia Association and has over four years of experience working at the intersection of data, software, and research. Jyoti’s work focuses on strengthening research software communities, improving research workflows, and supporting equitable participation in open science. She has led community initiatives, international collaborations, and research projects exploring the RSE landscape in Asia, including work supported through a Software Sustainability Institute Fellowship. With a background spanning statistics, software engineering, and community leadership, she is passionate about helping researchers develop sustainable software practices and build tools that can be shared across disciplines and regions.
From Internal Tool to Research Software: Preparing an R Package for CRAN
- Presenter: Jyoti Bhogal
- Authors: Bhogal, Jyoti*
Abstract: Many researchers develop useful scripts and tools for their own work, but relatively few evolve into sustainable research software that can be shared, reused, and maintained by others. This presentation shares the journey of transforming ‘generatervis’, an R package developed to simplify common data visualisation tasks, from a GitHub project into a CRAN-ready package. Making a package available through CRAN lowers barriers to discovery, installation, reuse, and long-term maintenance, helping research software reach communities beyond its original developers and institutions. The package development journey has been documented through public blog posts and an open-source repository, providing a transparent record of design decisions, implementation choices, and future development plans. The talk will explore the practical steps involved in improving software quality and preparing a package for wider adoption. Topics will include package design decisions, documentation, testing, continuous integration, dependency management, website generation with ‘pkgdown’, and meeting CRAN submission requirements. The presentation will also discuss challenges encountered during development and how these challenges influenced the package architecture and maintenance strategy. Beyond the technical process, this experience highlights the broader role of research software engineering in creating tools that are discoverable, reusable, and sustainable. The journey from a personal tool to a community-facing software package raises important questions about software quality, reproducibility, and long-term maintenance. By sharing lessons learned, mistakes made, and practical approaches used throughout the project, this presentation aims to support researchers, developers, and early-career RSEs who are interested in turning their own tools into sustainable research software. Attendees will leave with a clearer understanding of what is required to move beyond individual scripts and create software that can be reliably shared with the wider research community. Project Resources: 1) ‘generatervis’ GitHub Repository: https://github.com/Clinical-Informatics-Collaborative/generatervis 2) Development Blog Post 1: https://jyoti-bhogal.github.io/about-me/blog/2025-04-17_generatervis_blog_1/ 3) Development Blog Post 2: https://jyoti-bhogal.github.io/about-me/blog/2025-04-29_generatervis_blog_2/
Presenter Bio: Jyoti Bhogal (Bio details provided above).
The SSI Fellowship Without Borders: What Would Work, and What Would Not
- Presenters: Oscar Seip, Saranjeet Kaur, Aman Goel
- Authors: Seip, Oscar*
Abstract: For over a decade, the Software Sustainability Institute (SSI) Fellowship Programme has worked to improve and promote good computational practice across all research disciplines, while supporting the people who do this important work. It encourages its Fellows to be ambassadors for good practice in their own domains, fields, and areas of work, both in the UK and abroad. The SSI Fellowship model is increasingly recognised and adopted within the UK research software community as an effective way to build capacity and community. This raises a question: which elements of the model could work outside the UK, and how? And which elements could not? This panel begins by introducing the SSI Fellowship Programme and its history. In conversation with SSI staff and Fellows, it then explores which elements of the model might travel to different local contexts and which would not. Drawing on the Fellows’ lived experiences, the discussion explores how regional priorities, languages, knowledge systems, and resource realities across Asia, Australia, and the wider Global South would reshape such a programme, and how the SSI Fellowship itself would be enriched in return. Ultimately, the session shares lessons learned from the SSI Fellowship and offers a starting point for communities to consider what they might take from it, what they would do differently, and what the Fellowship has to learn from them. Audience members are invited to bring their own regional context and plans into this discussion.
Presenter Bio: Oscar Seip is a Research Community Manager at the Software Sustainability Institute (University of Manchester), where he manages the Fellowship Programme. Saranjeet Kaur is a 2023 SSI Fellow and founder of the RSE Asia community. She is also a Research Software Engineer at Imperial College London. Aman Goel is a 2023 SSI Fellow and SSI Research Software Training Manager, based at the University of Manchester.
The HelpfulBatBot: a RAG-based assistant for democratising Geodynamic modelling
- Presenters: Juan Carlos Graciosa, Louis Moresi
- Authors: Graciosa, Juan Carlos*; Moresi, Louis
Abstract: Understanding Earth processes, such as mountain formation, mantle convection, and groundwater flow, is inherently difficult due to the large range of length and time scales involved. While numerically modelling these phenomena have proven useful in gaining important insights, the associated modelling code may have relatively steep learning curves. One such code is Underworld3, a Python-based, open-source finite element, geophysical fluid dynamics modelling framework. Here, we present the HelpfulBatBot, a freely accessible web application using Retrieval-Augmented Generation (RAG) to assist users of Underworld3. This talk will highlight the following: 1.) structuring of the Underworld3 source code with Large Language Model (LLM)-readability in mind, 2.) the HelpfulBatBot RAG pipeline, with each component choice motivated by an Underworld3-specific challenge and evaluated through an ablation study, and 3.) a two-class evaluation framework for isolating RAG’s true value for domain-specific software.
Presenter Bio: Juan is a Research Officer at the Australian National University in the Research School of Earth Sciences. Funded by Auscope, he contributes to the maintenance and development of numerical models used by the Geodynamic modelling community. Prior to this, Juan completed his PhD at Monash University studying earthquakes with numerical models and explainable artificial intelligence.
Transcriptomics and protein-protein interaction analysis to reveal pathogenicity and antimicrobial resistance in children under 5 years with acute diarrhea
- Presenters: Linda Erlina, Fadilah Fadilah, Asmarinah Asmarinah, Badriul Hegar, Aryo Tedjo
- Authors: Erlina, Linda*; Fadilah, Fadilah; Asmarinah, Asmarinah; Hegar, Badriul; Tedjo, Aryo
Abstract: Diarrhea in children is often associated with dysbiosis and pathogenic microorganisms, thus requiring further molecular investigation. Transcriptomics analysis offers potential insights into biomarkers for the molecular pathogenic mechanisms of diarrhea and antibiotic resistance in children. This study aimed to identify differentially expressed genes (DEGs) associated with pathogenicity and antimicrobial resistance by analyzing publicly available transcriptomic data from two datasets (GSE69529 and GSE54070) in children with diarrhea. DEG analysis was performed using GEO2R on the GSE69529 and GSE54070 datasets. Several machine learning models (SVM, RF, NN, LR, XGBoost) were applied for cross-validation. Cytoscape and databases such as STRINGdb, IntAct, and CTD were used for protein interaction and upstream regulation analysis. Key differentially expressed genes (DEGs) included RYK, CREBL2, and ST14, which are associated with intestinal integrity, inflammation, and immune regulation. Seven resistance genes were identified as common genes across age groups. Pathogen mapping revealed Escherichia coli, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, and Salmonella typhimurium as the dominant pathogens. Resistance-associated genes such as mexA/acrA, arnA, tolC, and mexB/acrB are involved in the β-lactam and CAMP resistance pathways. Common resistance genes across age groups included tetO, tetQ, and SUL2. To validate the findings, Support Vector Machines and Neural Networks achieved the best classification performance among the machine learning models. Several genes identified among the DEGs associated with pathogenicity and antimicrobial resistance markers may serve as potential biomarkers to inform the development of diagnostic and therapeutic targets in pediatric diarrhea.
Presenter Bio: Currently, I’m a lecturer and junior researcher in the Department of Medical Chemistry and Bioinformatics Core Facilities, Indonesian Medical Education and Research Institute (IMERI), Faculty of Medicine, Universitas Indonesia, Indonesia. I have a research background in metagenomics (16s rRNA and shotgun metagenomic sequencing) and transcriptomics (Differentially Expressed Genes), and molecular simulation analysis (homology modeling, pharmacophore modeling, molecular docking, and molecular dynamic simulation), especially using Indonesian Herbal bioactive compounds for drug discovery and development. I am currently working toward completing my doctoral program in biomedical sciences at the Faculty of Medicine, Universitas Indonesia, on the topic of shotgun metagenomic sequencing and metabolomic analysis in children with acute diarrhea. I would be happy to collaborate in bioinformatics research. Here is my GitHub link: https://github.com/lindaerlina and email address [email protected].
Beyond Interoperability: Hands-On Federated ML for Research Data Curation Infrastructure
- Presenter: Arnab Mukherjee
- Authors: Arnab Mukherjee (SAFE Evidence Lab, Oklahoma State University Center for Health Sciences)
Abstract: Interoperability—the ability of systems to securely exchange, interpret, and use data without human intervention—has become a foundation of modern research data curation infrastructure. Yet interoperability alone does not establish whether data from different sites encode equivalent underlying phenomena. When clinical data exhibits a distributional shift across institutions, models trained at one site may fail silently at another. Confident but wrong is the worst possible failure mode, particularly when models are deployed across organizations. This is the foundational curation challenge this workshop addresses. Research software without borders requires not just that data can cross institutional lines, but that what the data means can be trusted on the other side. Site-level measurement heterogeneity determines whether data generated across diverse sources can be meaningfully compared, interpreted, and trusted as part of a shared research infrastructure. We complement interoperability-focused approaches by developing methods to identify and quantify site-level heterogeneity before data and models are integrated, thereby distinguishing true underlying variation from measurement-induced differences and establishing a more reliable foundation for curation in cross-site federated research. This workshop translates that framework into practice. Participants will engage in hands-on Python exercises using conformal prediction under simulated site-level heterogeneity, working directly with federated ML tools designed for research data curation pipelines. Attendees will leave with reusable code, a conceptual framework for distinguishing distributional shift from measurement-induced heterogeneity, and practical uncertainty quantification tools they can deploy within their own cross-institutional curation infrastructure.
Presenter Bio: Arnab Mukherjee is a Postdoctoral Researcher at the SAFE Evidence Lab at Oklahoma State University Center for Health Sciences focused on developing trustworthy, reproducible AI for health research. He earned a PhD in Computer Science with a specialization in Machine Learning from the University of Kansas. His work spans federated AI, machine learning systems, evidence generation, and large-scale biomedical data analysis. He leads the development of scalable ML and LLM pipelines, contributing to formal problem formulation, simulation design, evaluation frameworks, and automated evidence extraction. His research emphasizes open, interoperable, and reproducible machine learning infrastructure that enables transparent and reliable scientific discovery. Looking forward, he aims to advance federated AI systems and open-science platforms that improve the rigor, scalability, and real-world impact of biomedical and clinical research.
In the Beginning Was Documentation, and the Documentation Was Parsed, and the Documentation Begat the Code, and It Was Good
- Presenter: Elio Campitelli
- Authors: Elio Campitelli (Monash University)*
Abstract: The Climate Data Operator (CDO) is a command-line program that provides hundreds of subcommands, known as ‘operators’, for performing common operations in climate data processing. Due to its speed and efficiency, and because it works very easily with datasets larger than RAM, it is a very popular tool in the climate science community and one of the cornerstones of my scientific workflows. But, being a command-line tool, integrating CDO into a workflow using R or other programming language required me to concatenate text and calling system commands. This is very verbose and couldn’t take advantage of my IDE’s language features like auto-completion, syntax highlighting and quick access to documentation. So I wanted an R wrapper that would allow me to write normal R code that would be translated into CDO syntax, but with more than 700 operators, it looked like a monumental task. Fortunately, the design of CDO makes it possible to write an R wrapper automatically based on its documentation. This is how the rcdo package came about. What I actually developed is not the rcdo package itself, but a series of R scripts that build the package automatically. These scripts first use simple string manipulation to extract information about each operator from the CDO documentation: operator name, number of input files, number of output files and possible parameters. It also extracts the documentation itself: operator descriptions, parameter details and types. This creates an operator specification file that then it’s used to create a function for each operator and the corresponding help files. On top of that, rcdo provides several convenience features to integrate it even more within the R programming language. It manages temporary files automatically, has a garbage collection system that deletes output files when not longer used, and provides optional cache that avoids re-running potentially long operations. Automatically parsing the documentation had other side benefits. First, the operator specification can be used to create wrappers in other programming languages. I created the pycdo package, which supports the same features as its R sibling but using pythonic syntax. Second, I was able to analyse the documentation itself, finding errors and inconsistencies that are hard to find otherwise. I am now in contact with the CDO maintainers to help fix those issues, which will help all users, even the ones not using any wrapper.
Presenter Bio: Elio Campitelli is an atmospheric scientist studying Antarctic sea ice at Securing Antarctica’s Environmental Future in Monash University. They are an avid R user and have been amassing a collection of packages that help with their research for almost a decade.
Got a Data Challenge? There’s a Global Community for That: An Introduction to the Research Data Alliance (RDA)
- Presenter: Trish Radotic
- Authors: Trish Radotic (Research Data Alliance)*
Abstract: The Research Data Alliance (RDA) is a global, community-driven organisation bringing together researchers, technologists, data professionals, and infrastructure providers to solve research data challenges collaboratively across disciplines and borders. This lightning talk provides a practical introduction to the RDA ecosystem - how to join, find your community, contribute to existing Interest and Working Groups, and even start a new group around a shared data challenge or emerging topic. Attendees will gain a clear understanding of how open, international collaboration within the RDA can help transform local research data problems into globally supported community solutions.
Presenter Bio: Trish Radotic is the Regional Community Manager for Oceania and East Asia with the Australian Research Data Commons, based at the Curtin Institute for Data Science at Curtin University, Western Australia. She brings more than 20 years of experience across information technology, marketing, international business, and community engagement. Trish holds a Bachelor of Commerce in Accounting and Finance, an MBA in Marketing and International Business, and a Graduate Diploma in Software Programming. Through her work with the Research Data Alliance, she supports international open research communities and initiatives that advance the sharing and reuse of research data across disciplines, technologies, and borders. Her interests include open research, AI, and fostering global collaboration through community connection and knowledge sharing.
Research Software Without Borders: Building an Open Community Blueprint for Agentic AI in Research
- Presenter: Trish Radotic
- Authors: Trish Radotic (Research Data Alliance)*
Abstract: As agentic AI tools rapidly enter research workflows, research organisations face a fragmented landscape of proprietary platforms, shifting standards, and limited implementation guidance. In response, the Research Data Alliance (RDA) convened a global, community-driven consultation initiative to develop an open, technology-agnostic blueprint for responsible agentic AI implementation in research environments. This presentation shares how researchers, research software engineers, data specialists, librarians, and infrastructure providers collaborated across regions and disciplines to identify shared challenges, practical requirements, and guiding principles for the use of agentic AI in research. Rather than prescribing specific tools or vendors, the resulting blueprint provides a flexible, community-endorsed framework that organisations can adapt to their own institutional and technical contexts. Emerging themes include interoperability, transparency, provenance, human oversight, open standards, and FAIR-aligned design - reflecting broader tensions currently shaping the future of research software and data infrastructure globally. This presentation will provide an insight into the community consultation process, the structure and purpose of the blueprint, and opportunities to engage with an open international effort to support more transparent, sustainable, and community-driven AI adoption in research.
Presenter Bio: Trish Radotic is the Regional Community Manager for Oceania and East Asia with the Australian Research Data Commons, based at the Curtin Institute for Data Science at Curtin University, Western Australia. She brings more than 20 years of experience across information technology, marketing, international business, and community engagement. Trish holds a Bachelor of Commerce in Accounting and Finance, an MBA in Marketing and International Business, and a Graduate Diploma in Software Programming. Through her work with the Research Data Alliance, she supports international open research communities and initiatives that advance the sharing and reuse of research data across disciplines, technologies, and borders. Her interests include open research, AI, and fostering global collaboration through community connection and knowledge sharing.