Research
DeepMind Alumni Launch Faraday Agent to Automate Scientific Paper Reproduction
Inherent, a startup founded by former Google DeepMind researchers, has developed an AI agent called Faraday that the company says outperforms systems from Anthropic and OpenAI at reproducing experimental results from published scientific papers. The tool addresses a persistent challenge in research: validating whether published findings can be independently replicated.
By Michael G ·

Reproducing scientific results remains one of the field's most labor-intensive bottlenecks. When researchers publish findings, validating those results often requires months of work: reverse-engineering experimental setups, locating datasets that may be proprietary or lost to time, reconstructing software environments from sparse documentation, and debugging code written by authors who may no longer maintain it. Inherent, founded by former Google DeepMind researchers, says its Faraday agent can automate much of this work. According to a TechCrunch report, the company claims Faraday outperformed comparable systems from Anthropic and OpenAI in reproducing results from scientific papers across multiple domains.
The startup's focus reflects a real crisis in academic publishing. Studies over the past decade have documented that many published results cannot be reproduced by independent teams. Some failures stem from incomplete documentation. Others result from subtle dependencies between code, data, and compute environments that authors never anticipated would matter. Still others reveal negative results: experiments that worked under specific conditions but fail when parameters shift. Faraday targets this problem by treating paper reproduction as a structured engineering task. The agent reads published papers, extracts methodological details, searches for associated code and datasets, reconstructs environments, and attempts to run experiments end-to-end. When it succeeds, researchers gain confidence that results are genuine. When it fails, the specific point of failure often reveals whether the issue stems from missing data, undocumented environment dependencies, or actual problems with the original work. TechCrunch report documents the reporting behind this account.
Distinguishing code reproduction from scientific validation matters here. Faraday can execute code and replicate numerical outputs. That is different from independently validating a scientific conclusion. A paper's code may run identically on a second machine yet still draw conclusions not fully supported by the data, or may fail to account for confounding variables that an independent replication effort would catch. What Faraday addresses is the engineering layer: the practical, time-consuming work of getting published code to run in a new context. That layer is substantial enough that automating it creates real value for research teams and publishing platforms alike.
The Replication Bottleneck
Hidden data represents one of the hardest obstacles to replication. Papers sometimes rely on proprietary datasets, patient records restricted by privacy law, or internal corporate data that cannot be shared publicly. Authors may have trained on arXiv papers or other open sources, but exact dataset versions drift over time as repositories update or remove files. Code may reference local file paths or hardcoded credentials removed before publication. Faraday cannot conjure missing data, but it can identify what is missing and flag it explicitly. That diagnostic step alone saves researchers weeks of dead-end debugging. The agent learns which gaps are fatal to reproduction and which can be worked around. It also distinguishes between cases where authors have genuinely provided insufficient detail versus cases where the necessary information exists in supplementary materials, GitHub repositories, or author-published blogs that the paper itself does not cite. Inherent offers useful technical background for evaluating the claim.

Environment reconstruction also looms large. A paper published five years ago may have depended on specific versions of machine learning libraries, Python releases, GPU drivers, or operating systems. Those versions often no longer work together. Dependency trees have shifted. Security patches have broken old APIs. Faraday must navigate this archaeological layer by reasoning about version compatibility, identifying likely substitution paths, and testing whether modern libraries can replicate the original numerical outputs. Some experiments are sensitive to hardware: floating-point rounding, CPU cache behavior, GPU architecture differences, or random-seed initialization. Faraday needs to account for these physical constraints and report when results diverge due to hardware rather than method. The operational tradeoff is also reflected in Google DeepMind.
The company's benchmark claims come from Inherent's own testing. The specific papers, metrics, and competing systems involved remain proprietary. That does not invalidate the claims, but it means the research community cannot independently verify the numbers without access to Inherent's test set and methodology. Third-party evaluation would require researchers to run Faraday, other agents from Anthropic and OpenAI, and comparison systems on a shared set of papers, measuring success by consistent criteria. That work has not been published as of August 2026. Inherent has created the tool and demonstrated it to early users and press, which supports the credibility of the basic capability. Validating the magnitude of the performance gap requires external replication. The company has published a website outlining the product, but comprehensive technical documentation of Faraday's architecture, training approach, and limitations remains limited. For broader context, arXiv outlines the relevant standard or institution.
Implications for Publishing and Research

If Faraday performs reliably at scale, the tool could reshape how journals handle reproducibility. Some publishers have begun requiring authors to submit code and data with papers. Others run basic checks on submitted code before acceptance. An automated agent that can reproduce published results at the point of publication, or shortly after, would offer a new enforcement mechanism. Journals could use replication success as a metric in peer review or post-publication evaluation. Researchers citing a paper could query whether Faraday has successfully reproduced it. Funding agencies could require that grant recipients demonstrate reproducibility of their published results using such a tool. These moves raise their own questions: How should journals weight reproduction success against novelty? Should negative replication results be published? Should papers that fail replication be retracted, or is partial replication valuable in itself? The answers depend on the field, the stakes, and the reason for failure. A machine learning paper whose code fails to run due to a missing dataset may still contain valuable ideas. A clinical trial result that cannot be replicated may pose direct safety risks. Faraday can automate the technical work without resolving these policy questions.
The broader context here is the growing complexity of machine learning research itself. Papers in the field now routinely depend on large codebases, proprietary datasets, distributed compute infrastructure, and commercial APIs. Replicating a deep learning result means controlling not just code and data but also GPU availability, training time, and random initialization effects that can shift results significantly. Faraday must reason about all of these factors. It must also handle papers that combine machine learning with other domains: robotics papers that depend on physical hardware specifications, neuroscience papers that depend on proprietary brain imaging data, or drug discovery papers that depend on computational chemistry libraries. No single agent can handle all cases perfectly. Faraday's advantage, according to Inherent, is that it outperforms existing general-purpose AI systems on this specific task. That may reflect targeted training, better prompting strategies, or domain-specific reasoning built into the agent.
Independent replication of Faraday's own performance will ultimately determine its impact. The research community should treat company claims about benchmark performance with appropriate caution, especially when the benchmarks are not yet public. What matters is whether Faraday succeeds when deployed by actual researchers on papers they care about, across enough diverse fields and enough papers to generate reliable data. Early adopters will be critical here. As more researchers use Faraday and report results, patterns will emerge about which types of papers it handles well and which ones remain intractable. That user feedback will shape both the tool and the conversation around how AI can accelerate reproducible research. For now, Faraday represents a meaningful engineering advance toward automating one of research's most frustrating tasks. Whether it becomes a standard layer in scientific publishing depends on validation, accessibility, and sustained development beyond the initial launch.
Topics: AI research, scientific reproducibility, machine learning, DeepMind, automation