中文 Apply now

01PhAI Labs · Toward real scientific discovery

DFM Scientist Collaboration Program

Bring frontier questions into the loop of intelligence

Connect real scientific questions, data and experimental feedback with next-generation AI—and explore how AI can move from assisting research to participating in discovery.

Apply now Explore DFM

The first cohort prioritizes life science and biomedicine, while remaining open to chemistry, materials and other frontier fields.

Explore the program

02Why this program

Discovery needs more
than a faster tool

AI can already retrieve literature, generate code, process data and assist with parts of the research workflow. But discovery often begins with a question that is not yet well defined—and advances through hypotheses, evidence, tools, experiments and continuous revision.

That process cannot happen only on static benchmarks or closed tasks. AI must enter real research environments, confront incomplete data and constraints, and learn from experiments and reality.

  1. 01

    Start with real questions

    Begin with the unknowns, bottlenecks and anomalies scientists actually face.

  2. 02

    Verify through real feedback

    Subject model outputs to data, experiments, expert judgement and the physical world.

  3. 03

    Improve through collaboration

    Turn questions, tools, data and verifiers into reusable infrastructure for the next discovery.

03What is DFM

From completing tasks
to participating in discovery

DFM (Discovery Foundation Models) is neither a single model nor a software framework that forces every system through one pipeline. It is the research paradigm, capability definition and system organisation that PhAI Labs proposes around Discovery Intelligence: through formal definitions, process diagrams, system instances and real scientific cases, it turns “discovery” from a vague notion into a capability of model systems that can be learned, executed and evaluated.

The DFM Scientist Collaboration Program brings real scientific problems, data and verification conditions into this paradigm, so systems form their judgement in real research.

Discovery Foundation Models technical diagram
Recursive discovery loop · Partially understood world → Valuable unknown → Researchable problem → Explanation & intervention → Validated new knowledge

Three independent research and engineering directions

ScienceBuddy, ScienceIDE and JEPA-Anything are three fully independent modules and research directions: independent in product form, usage, papers and implementation, with no required calls or dependencies between them. Looking ahead, the three could form a composable research path: ScienceBuddy collects feedback and research trajectories from the interaction between scientists and agents; ScienceIDE organises that data and experience into reusable scientific environments for task execution, reinforcement learning and model training; JEPA-Anything explores unified state prediction and simulation across multiple scientific environments and interventions. That is a possible future combination, not an integration that exists today, and it does not change their independent standing.

16 Sept

ScienceBuddy · AI research partner

An interactive Scientific Agent, exploring RSI in RSI in real research interaction

It focuses on an interactive form of Science RSI: a user poses a scientific question, the agent calls tools and answers; the scientist's follow-ups, edits, acceptances, rejections and re-runs become feedback that drives the research agent's workspace and harness to keep evolving.

Open ScienceBuddyReleasing 16 Sept
17 Sept

ScienceIDE · science integrated development platform

A platform integrating scientific environments, code, tools and training pipelines

It turns real scientific code, research procedures, task goals and verification criteria into scientific environments that run, measure and reproduce, so that researchers can do agent execution, model training and reinforcement-learning research in real scientific settings.

ScienceIDE project pageReleasing 17 Sept
18 Sept

JEPA-Anything · cross-domain unified science world model

A unified science world model for cross-domain state prediction, intervention simulation and scientific diagnosis

It explores how to learn, across vision, biology, the clinic, molecules, control and physics, latent representations that predict changes of world state, and to offer one interface for intervention prediction, state synthesis and scientific diagnosis.

JEPA-Anything paperReleasing 18 Sept

04What we look for

What questions are worth exploring together?

We are most interested in questions with real scientific value, a consequential unknown, and a path to data or experimental feedback—not tasks already fully specified and solved by an existing workflow.

  1. 01

    A consequential unknown

    A question that could explain an unresolved phenomenon, reveal a new mechanism or pattern, or lead to a new way of doing research.

  2. 02

    A real bottleneck

    Current theory, models, data analysis or experiments do not provide a satisfying answer—and the failure itself is informative.

  3. 03

    A workable starting point

    Some combination of data, literature, tools, experimental conditions or domain knowledge makes a first exploration possible.

  4. 04

    A verifiable signal

    A hypothesis can be supported, corrected or rejected through data, experiments, expert judgement or the real world.

Questions to begin with

  1. 01What phenomenon in your research remains unexplained?
  2. 02Why do current methods fail to resolve it?
  3. 03What data, tools or experimental conditions already exist?
  4. 04What outcome would count as meaningful progress?

05First cohort & who should join

Bring judgement, data and verification
into one discovery loop

The program is open to researchers in universities, institutes, laboratories and industrial R&D teams. Questions may be at the stage of early conception, data accumulation, methodological bottleneck or experimental validation.

Initial focus: life science and biomedicine

These fields combine rich multimodal data, specialized tools and experimental verification with a large number of open, consequential scientific questions.

  • Omics & single-cell analysis
  • Disease mechanisms & biomarkers
  • Molecules, proteins & biological systems
  • Experiment design & testable hypotheses

The collaborators we hope to explore with

  1. 01

    Scientists and research groups with a question, wanting AI built in

    You know which questions matter and where current methods stop, and you want foundation models, research agents or automated experimentation wired into how you already work. We start from the specific bottleneck, work out where AI actually helps, then design and test the approach together.

  2. 02

    Labs working mainly through computation, modelling and simulation

    You have specialist data, research code, computational models or simulation environments, and want AI to extend how you analyse and explore. We look at how AI can run verifiable research tasks inside them.

  3. 03

    Wet labs looking to pair experiments with AI

    You run experiments, collect observations and test hypotheses, and want to know whether AI can help design the experiment, read the result or choose the next direction. We close the loop from question to evidence together.

  4. 04

    Early-career PIs, postdocs and PhD students trying a new way of working

    Researchers who want AI deep in the research process and are willing to work out a new paradigm with us. If you have a question you keep returning to, relevant groundwork, or a direction you want to explore together, tell us about it.

  5. 05

    Industrial R&D teams facing hard development problems

    You are working on the science behind a product or a process, with business context, data or test conditions already in hand, but current methods cannot explain the key behaviour, or it takes too many trials to find something that works. We identify where AI fits and measure progress by computation, simulation or experiment, against your real requirements.

06Collaboration

Not a request queue—
a research partnership

What PhAI Labs brings

AI research & engineering
Support across foundation models, agents, reinforcement learning, tool use and scientific intelligence.
Models, tools & compute
Project-specific access to relevant models, data tools and compute resources.
Problem formulation & validation
Turn an open question into an explorable and verifiable research path.
Research outputs & communication
Explore joint research, publications or open-source releases for substantive outcomes.
Long-term co-research & funding support
Projects that demonstrate early validity, aligned goals and sustained research value may progress into longer-term joint research with corresponding funding support. The form of support will be determined together as the project develops.

What scientists bring

  • Explain why the question matters and where current methods fail
  • Contribute essential domain knowledge and existing research context
  • Provide authorized access to data, tools or experimental conditions
  • Judge analyses and hypotheses, and provide verification feedback

Data & intellectual propertyBefore substantive work begins, the parties will document data scope, confidentiality, publication, authorship and IP arrangements for the specific project.

  • Use data only within an explicitly authorized scope
  • Clarify separately whether data may improve or train models
  • Protect unpublished data and results under agreed terms
  • Determine papers, code, models and patents through contribution and written agreement

07Exploration cases

Begin with real scientific questions

The first seed users are bringing real research questions into testing. With their permission, we will share representative exploration processes, intermediate outcomes and lessons.

Rather than display only a final answer, we want to show how a problem was framed, where AI participated, how scientists judged the output, and how the research moved forward.

Case 01

GALILEO

Autonomous therapeutic discovery by an embodied AI scientist

Generalizable Agentic Laboratory Intelligence for Learning, Experimentation, and Optimization

GALILEO is the DFM idea made concrete in the life sciences: AI is no longer a tool that completes a single task but enters a real research environment, starts from a real question, keeps learning from experimental feedback and takes part in the whole process of discovery. The work is available as a bioRxiv preprint (June 2026).

Research question

Most AI systems in biomedicine stop at prediction: predicting targets, predicting molecular properties, without reaching real experimental intervention. GALILEO asks whether AI can move past prediction and close the loop autonomously in a dynamic membrane system, from candidate peptide discovery to wet-lab validation, and find therapeutic targets and molecules that intervene in tumour immunity.

What was already in place
  • A clinically informed peptide prior library (CPP) giving candidates a traceable starting point
  • A robotic solid-phase peptide synthesis platform that synthesises AI-designed molecules automatically
  • Multimodal phenotyping across membrane current, metabolic flux and T-cell function
  • Public omics data from TCGA, CPTAC, GEO and GTEx, plus a wet-lab validation team
How AI takes part

GALILEO runs on an OTAS reasoning loop (Observation, Thought, Action, Summary) and executes autonomously:

  • Retrieving candidates from the peptide prior library and editing sequences through auditable local operations
  • Ranking target branches on its own, prioritising the LRRC8C and SLC25A1 pathways
  • Updating target beliefs, sequence strategy, assay design and mechanistic hypotheses from wet-lab feedback
  • Keeping the whole process traceable and auditable, so scientists can inspect and intervene at any point
Stage of progress
  • The loop from prediction to intervention has been validated and released as a bioRxiv preprint
  • Generated peptides blocked LRRC8C current, perturbed osmotic and redox homeostasis and promoted tumour-dependent T-cell activation
  • Peptides on the SLC25A1 branch disrupted citrate export metabolism, reduced extracellular acidosis and strengthened CD8+ T-cell function
  • Publication and translation are under way; case details will be disclosed progressively once researchers authorise it

The DFM viewGALILEO shows what DFM argues for: moving from completing research tasks to taking part in discovery. AI enters a real research environment, meets the questions scientists actually care about, works with incomplete data and real constraints, and keeps learning from experiments and reality, so each exploration becomes the starting point of the next discovery.

08Collaboration process

From application to joint research

  1. 01

    Apply

    Describe the scientific question, current progress, key bottlenecks, and available data or verification.

  2. 02

    Evaluate & select

    Assess scientific value, the role for AI, verification feasibility and mutual resource fit.

  3. 03

    Initial conversation

    Clarify the scientific context, collaboration goals and current research foundation together.

  4. 04

    Co-design

    Define goals, workflow, contributions, data use and stage-gate criteria together.

  5. 05

    Collaborate

    Move into early validation or deeper joint research and iterate through real feedback.

09Frequently asked questions

What you may
want to know

01What questions are the strongest fit for the first cohort?

We prioritize consequential scientific unknowns where AI can make a substantive contribution and where data, experiments or expert feedback can support verification. The question need not be fully formalized, but its importance and a viable starting point should be clear.

02Can I apply with an early-stage question rather than a complete proposal?

Yes. Identifying a valuable unknown and turning it into a researchable problem is itself part of the DFM agenda. Share the observation, the limits of current methods, and the evidence or conditions already available.

03How large is an initial collaboration?

We prefer to begin with a bounded subproblem and test the problem formulation, data conditions, role for AI and verification path before considering deeper joint research. Scope and pace are co-designed.

04How does PhAI Labs evaluate fit?

We consider scientific importance, whether DFM or the Research Agent can create genuine research leverage, the feasibility of data and verification, and whether both sides can commit the right research and engineering resources.

05What outcomes can a collaboration produce?

Outcomes may include a validated hypothesis, reusable research workflow, tool or data asset, model or Agent improvements, and where appropriate, papers, open-source releases or other joint outputs.

06What should not be included in the first application?

Do not upload patient-identifiable, clinically sensitive, third-party restricted or otherwise unauthorized or confidential material. Appropriate data exchange and protection can be discussed later.

10Apply now

Start with a question not yet fully understood

You do not need a complete project proposal. Tell us what you want to understand, why it matters, where the current work is blocked, and what data, tools or verification conditions already exist.

Apply now Feishu form

Please do not upload patient-identifiable, clinically sensitive, third-party restricted or otherwise unauthorized material in the initial submission.