Capabilities

The problems I am good at.

Stated as problems rather than methods — because nobody has a shortage of methods. Each entry is a situation I have worked in, what I do about it, and what comes out the other end.

The same toolkit shows up in two settings: product and platform work in health-tech, and research questions in labs and clinics. The problems differ more than the methods do.

Setting one

Health-tech and product work

Data, a hypothesis and a deadline. What is usually missing is someone who can establish whether the signal is real before a product gets built on top of it.

Problem 01

You have a cohort and no idea where the signal is

Thousands of analytes, hundreds of thousands of rows, a dozen plausible endpoints, and an analysis that keeps producing findings that evaporate on the next dataset.

Approach
Nail the phenotype and endpoint definitions first. Screen across all analytes with proper multiple-testing control, model non-linear dose–response rather than assuming linearity, then confirm with time-to-event models and pathway-level interpretation.
Outcome
A ranked, FDR-controlled biomarker panel with effect sizes and confidence intervals, the biology it implicates, and an honest account of what the cohort cannot tell you.

Problem 02

Your model looks excellent in the paper and fails in the clinic

High AUC on the development set, disappointing performance everywhere else. Usually this is leakage, distribution shift, or a model that ranks well but whose probabilities mean nothing.

Approach
A structured audit: data-flow review for leakage, patient-level split integrity, recalibration, external or temporal validation, subgroup performance, and decision-curve analysis at the thresholds you actually use.
Outcome
A validation report naming the specific failure modes, a recalibrated model where recalibration is enough, and a clear statement of the population the model is licensed for.

Problem 03

You need to know whether the study is worth running

An expensive trial, a costly assay, or a new measurement — and no way to know in advance which arm, endpoint or sampling schedule gives you a chance of detecting anything.

Approach
Build a mechanistic model of the process, generate a virtual cohort, and run the study in silico under realistic noise and dropout. Vary the design; see which versions can actually resolve the effect.
Outcome
A design recommendation backed by simulation — sample size, timing, which measurement carries the information, and which planned measurements are not worth their cost.

Problem 04

You have imaging and outcomes but no pipeline

MRI, histology or microscopy sitting in a folder, an outcome column, and a cohort far too small for the deep-learning approach everyone suggests.

Approach
Reproducible segmentation and quantitative feature extraction, feature selection that is stable under resampling, and models sized to the cohort you have rather than the one you wish you had.
Outcome
A running image-to-prediction pipeline, the stable feature set, and validation that survives the small-n regime.

Problem 05

Your analysis cannot be re-run

The results exist, the person who produced them has moved on, and nobody can reproduce the figure — which becomes urgent exactly when a regulator, a referee or an investor asks.

Approach
Rebuild the analysis as a version-controlled, environment-pinned pipeline with provenance from raw extract to final figure, documented decisions, and tests on the steps that matter.
Outcome
A pipeline anyone on the team can re-run to identical numbers, plus documentation of every judgement call the original analysis made silently.

Problem 06

Your data exists, but not in a form anything can analyse

A biobank, a registry or a cohort whose records were never collected with analysis in mind, so every new question begins by assembling the data by hand.

Approach
Design a data model that fits how the data is actually collected, build the path that digitises it, and put the analysis behind an application the team runs itself. I built this for the Perakakis Lab's in-house biobank.
Outcome
A structured dataset and the software that maintains it — plus analyses the team can run directly, instead of rebuilding them by hand each time.

Problem 07

You want to use LLMs without inheriting their failure modes

Everyone wants AI in the product. Nobody wants a confident, fluent, wrong answer attached to a clinical or scientific claim.

Approach
Scope the task to something checkable, ground it in retrieval over sources you control, structure the output so it can be validated automatically, and build an evaluation set before shipping. Multi-provider routing so one vendor's outage or price change is not an outage of yours.
Outcome
A working pipeline with an evaluation harness — measured accuracy on your own task, cited sources, and a clear boundary around what the system is allowed to assert.

Setting two

Research and clinical questions

The biology, the patients and the question are already there. What is missing is turning it into something a model can answer — and staying with it through the referee reports.

Question 01

“We have a mechanism but no way to test it”

A hypothesis about how cells, tumours or immune populations interact that no experiment can observe directly at the relevant scale.

What I do
Formalise the mechanism as an agent-based, hybrid or reaction–diffusion model, calibrate it to whatever data exists, and identify the observable consequences that would distinguish your hypothesis from its competitors.
Outcome
A model you can interrogate, a set of testable predictions, and usually a figure that reframes the paper.

Question 02

“We have omics data and a statistician-shaped hole”

Proteomic, metabolomic or transcriptomic data alongside clinical variables, and an analysis plan that stops at “then we do machine learning”.

What I do
Integration across omic and clinical layers, cross-sectional and longitudinal; batch and covariate handling; multiple-testing strategy; and interpretation that lands at pathway level rather than a list of accession numbers.
Outcome
Analysis, figures and a Methods section written to survive review — plus the code, so the revision does not restart from zero.

Question 03

“Is this study designed well enough to be worth doing?”

Grant deadline approaching, endpoints not settled, and a power calculation nobody quite believes.

What I do
Endpoint and phenotype definition, realistic power analysis by simulation rather than formula, sampling schedule design, and an analysis plan written before the data exists.
Outcome
A pre-specified analysis plan and a defensible design section — the parts reviewers attack first.

Question 04

“We need modelling in the proposal, and someone to run it”

A consortium proposal that needs a credible computational work package, written by somebody who will actually deliver it.

What I do
Draft the modelling and data-analysis work package, scope it to something deliverable, and join the consortium to execute it. I also co-supervise doctoral students on the modelling side.
Outcome
A work package that reads as feasible because it is, and a named person attached to it.

What I will not do is hand you a number I do not believe. If the data cannot answer the question, the most valuable thing I can deliver is that finding — early, and in writing.

How I work with people on this →

Get in touch

Have data and a question that needs answering?

I work with researchers, clinicians, health-tech teams and bioinformaticians at the interface of biology, data science and personalised medicine — from a one-hour sanity check on a study design to a full modelling collaboration.

Details

Email
pejman.shojaee@tu-dresden.de
ORCID
0000-0003-3298-3315
Based in
Dresden, Germany
Languages
English (C1) · German (B2) · Persian (native)