中文 Apply now

PhAI Labs · Tech

Technology for scientific discovery

PhAI Labs' technical work turns on one question: how can AI systems take part in real scientific discovery. DFM is the research paradigm we propose for it; ScienceBuddy, ScienceIDE and JEPA-Anything are three independent research and engineering directions, working from interactive feedback, reusable scientific environments and world-state prediction respectively.

  1. Scientific question
  2. Model training
  3. Tool use
  4. Real experiments
  5. External evidence
  6. Capability growth

01Our technical direction

AI can already retrieve literature, generate code, process data and assist with parts of the research workflow. But discovery usually begins with a question that is not yet well defined, and advances through hypotheses, evidence, tools, experiments and continuous revision. That process cannot happen only on static benchmarks or closed tasks.

Intelligence has run out of internet. Published data is static and already trained on. What the frontier is short of is scientific data with a verified outcome attached: a hypothesis confirmed or rejected, an experiment that worked or failed. That data cannot be scraped. It has to be produced, inside real research.

So our direction is scientific intelligence itself: systems that keep learning, verifying and improving through real tasks and environmental feedback. Data trains the model; the model designs the next experiment; what comes back, including every failure, is the signal for the next round. This is why PhAI Labs works on foundation models, reinforcement learning, agents and scientific intelligence, and it is where DFM starts.

02DFM · Discovery Foundation Models

A research paradigm around Discovery Intelligence

DFM (Discovery Foundation Models) is neither a single model nor a software framework that forces every system through one pipeline. It is the research paradigm, capability definition and system organisation that PhAI Labs proposes around Discovery Intelligence: through formal definitions, process diagrams, system instances and real scientific cases, it turns “discovery” from a vague notion into a capability of model systems that can be learned, executed and evaluated.

DFM technical diagram: three stages of intelligence, current foundation models versus discovery foundation models, and the recursive discovery loop
DFM technical diagram · click to open full size

DFM is not a model, and not a software framework that forces every system through one pipeline. It is a paradigm, a capability definition and a way of organising systems: which capabilities the discovery process needs, how they connect, and where external evidence has to enter.

In this process, the scientific question comes first. Starting from a partially understood world, the system identifies a valuable unknown, turns it into a researchable problem, and builds hypotheses around it. A hypothesis does not stay on paper. The system calls research tools, data and experimental environments to act on it, gathers external evidence from analysis, experiments and expert judgement, and uses that evidence to confirm, revise or reject its own position.

Model training sits inside this process rather than in front of it. Every tool call and every experimental result is both a test of the model’s judgement and a training signal with a real outcome attached. Validated knowledge, reusable questions, tools and verifiers accumulate, and become the starting point for the next discovery.

Around this paradigm and capability definition, the research unfolds through system instances, independent papers, engineering modules and real scientific collaborations, released progressively. They do not add up to one fixed implementation; they explore, from different directions, how Discovery Intelligence becomes operational.

DFM takes an AI through a whole turn of discovery: find a question worth asking, test it against real evidence, keep what worked for next time.

Move the cursor over the drawing

The six capabilities DFM is about

  1. 01

    Identify valuable unknowns

    Start from unexplained phenomena, places where methods fail, and anomalies in data, and judge which unknowns are worth pursuing.

  2. 02

    Turn unknowns into researchable questions

    Narrow a vague unknown into a bounded question with a way in and a way to check it, and state what would count as progress.

  3. 03

    Build and revise hypotheses

    Propose hypotheses around the question, organize the existing evidence, and revise or drop a hypothesis when new evidence arrives.

  4. 04

    Call tools, data and experimental environments

    Take the hypothesis off the page: call literature, code, analysis tools and experimental environments, and turn judgement into actions that can be executed.

  5. 05

    Validate and update against external evidence

    Subject the system’s outputs to data, experiments, expert judgement and the physical world, and update its position according to what comes back.

  6. 06

    Accumulate reusable research structure

    Turn validated questions, tools, data and verifiers into reusable structure, so that each exploration becomes the starting point for the next discovery.

The DFM research-map repository carries the conceptual structure of the Discovery Loop, research directions, the open-source plan, demos and how to collaborate. Its directories and direction markers reserve future scope in public; they do not mean the corresponding implementations are finished or released.

03Three independent directions

ScienceBuddy, ScienceIDE and JEPA-Anything are three fully independent modules and research directions: independent in product form, usage, papers and implementation, with no required calls or dependencies between them.

04Independent direction·16 Sept

ScienceBuddyAI research partner

An interactive Scientific Agent: recursive self-improvement in real research interaction, exploring Recursive-in-Recursive Self-Improvement (RSI in RSI).

ScienceBuddy is an independent research and engineering exploration built around interaction with scientists, research feedback and the continual improvement of agents, in the Discovery Intelligence direction DFM is about.

It focuses on an interactive form of Science RSI: a user poses a scientific question, the agent calls tools and answers; the scientist's follow-ups, edits, acceptances, rejections and re-runs become feedback that drives the research agent's workspace and harness to keep evolving.

What ScienceBuddy produces is not only one-off answers but the feedback, research trajectories, tool calls and task states of the interaction. In future these can be organised into reusable scientific tasks and environments for reinforcement learning or other model training.

An early product / research preview: the capabilities opened, demonstrations and entry points follow the actual version. The product entry opens on 16 September.

A scientist and an AI partner talk it through; every follow-up, edit and acceptance is written down, and the AI's workspace gets better round by round.

Move the cursor over the drawing

It takes in and structures

  • Scientific questions
  • Data needs
  • Research materials
  • Literature and code
  • Scientific tool calls
  • Hypothesis formation and testing
  • What the following research workflow needs

05Independent direction·17 Sept

ScienceIDEScience integrated development platform

Infrastructure for scientific experience: real scientific code, tasks, environments and reinforcement-learning training.

ScienceIDE is an independent infrastructure exploration for real scientific codebases, research tasks and verifiable scientific experience.

It turns real scientific code, research procedures, task goals and verification criteria into scientific environments that run, measure and reproduce, so that researchers can do agent execution, model training and reinforcement-learning research in real scientific settings.

Its value is in organising real scientific software, domain judgement, task goals and verifiable feedback into one executable environment, an experiential base for research agents, model training, reinforcement learning and discovery research. ScienceIDE does not require ScienceBuddy to be used, and does not depend on JEPA-Anything to stand.

ScienceIDE is open for task proposals. Whether the project page, submission page and repository go public on 17 September follows final readiness.

Turn a real scientific codebase into a row of workbenches that run, get scored and can be repeated, so an AI can work on real scientific tasks and be checked.

Move the cursor over the drawing

Current task forms

  • Accelerate
  • Discover
  • Repair
  • Reproduce
  • Integrate

06Independent direction·18 Sept

JEPA-AnythingCross-domain unified science world model

A cross-domain unified science world model: cross-domain state prediction, intervention simulation and scientific diagnosis.

JEPA-Anything is an independent model-method study of latent world models and orthogonal predictive factorisation.

It explores how to learn, across vision, biology, the clinic, molecules, control and physics, latent representations that predict changes of world state, and to offer one interface for intervention prediction, state synthesis and scientific diagnosis.

In the research vision, JEPA-Anything can be read as a unified predictive model or science world simulator over many scientific environments: when a researcher proposes an action or intervention, the model predicts the likely state changes, outcomes and risks in advance, helping screen out ineffective or costly plans. That is a direction to explore: today it does not simulate every scientific environment, and it is not connected to ScienceBuddy or ScienceIDE.

Paper, code, model weights and demo links go public on 18 September within the confirmed scope.

Predict before acting: if we do this, what does the world look like, and which option is worth taking.

Move the cursor over the drawing

Research questions

  • How to represent the many predictable factors of a complex system
  • How to predict the state changes different interventions bring
  • How to ground experiment selection and scientific diagnosis in prediction
  • How to reuse a latent world-state interface across domains

07A composable research path

A possible future path

Looking ahead, the three could form a composable research path: ScienceBuddy collects feedback and research trajectories from the interaction between scientists and agents; ScienceIDE organises that data and experience into reusable scientific environments for task execution, reinforcement learning and model training; JEPA-Anything explores unified state prediction and simulation across multiple scientific environments and interventions. That is a possible future combination, not an integration that exists today, and it does not change their independent standing.

  1. ScienceBuddy keeps interacting with scientists, producing feedback, task states and research trajectories
  2. ScienceIDE organises some of that experience into reusable scientific environments for executing, training and evaluating models
  3. JEPA-Anything explores unified state prediction over those environments and candidate interventions, so researchers can screen plans before acting
  4. Predictions and real experimental results return to the research process as new evidence

The path expresses the potential synergy of three independent works. It is not a completed integration, and it does not mean the three must merge into one system.

One module at a time

DFM is a paradigm, not a finished model. The research around it unfolds through system instances, independent papers, engineering modules and real scientific collaborations: on 15 September the DFM technical report and research-map repository go public; on 16, 17 and 18 September ScienceBuddy, ScienceIDE and JEPA-Anything are released in turn. Each project's links, code and product entries open on their day; anything that exists earlier is marked “coming soon”.

Four-day release calendar

  1. The DFM paradigm and the DFM Scientist Collaboration Program
  2. ScienceBuddy · AI research partner
  3. ScienceIDE · science integrated development platform
  4. JEPA-Anything · cross-domain unified science world model

All news →