Home / Resources

Resources

133 entries

Workbenches 64

Local-first desktop workbench wiring agents, notebooks, runs, figures and review into one auditable provenance trail.

1.6k

Galaxy

Nucleic Acids Res

Galaxy for accessible, reproducible, and collaborative data analyses: 2026 update

Plans and implements ML engineering work with arXiv integration and code retrieval.

1.6k

Asta

2025

AI2's science agent family, reproducible and benchmarkable against a rigorous multi-task research suite.

Autonomous deep research over web and local documents with any provider, emitting a cited report.

Long-horizon agent harness that researches, writes code, and produces artifacts.

Fully local, encrypted research agent over arXiv, PubMed, and your own private document collection.

9.1k

Hierarchical planner plus specialist agents for deep research and general task execution.

3.5k

Open research assistant combining search, code execution, link resolution and information expansion.

108

Write Python protocols and execute them on physical Flex and OT-2 liquid-handling robots.

Orchestrates beamline and laboratory experiments plus data acquisition, in production at NSLS-II.

Agent skills for topic exploration, literature survey, experiments, paper writing and integrity audit.

215

Multi-agent simulation of science-of-science dynamics over real publication data.

143

Robust Molecular Structure Recognition with Image-to-Graph Generation

334

CAMEL

2023

Communicative Agents for "Mind" Exploration of Large Language Model Society

OWL

2025

Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

Minimal library for code-writing agents, shipping the reference Open Deep Research implementation.

Gives an agent a real single-GPU nanochat training setup and lets it modify code, train, evaluate, keep or discard.

95k

Markdown-only skill pack for autonomous ML research providing cross-model review loops, idea discovery and experiment automation.

16k

Semi-automated research assistant spanning ideation, coding, experiments, writing and publication across Claude Code, Codex, Kimi and OpenCode.

5.4k

Self-evolving research colleague with 285 skills across 28 disciplines and persistent memory over literature and databases.

MLGym

2025

A New Framework and Benchmark for Advancing AI Research Agents

CMBAgent

ICML Workshop2025

Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery

LLM-SR

ICLR (Oral)2025

Scientific Equation Discovery via Programming with Large Language Models

AI Hilbert

Nature Communications2024

Evolving scientific discovery by unifying data and background knowledge with AI Hilbert

CLI and leaderboard that autonomously optimizes existing research codebases, publishing results only when an internal ledger confirms improvement.

MDCrow

MLST2026

Automating Molecular Dynamics Workflows with Large Language Models

245

LLaMP

EMNLP Main2025

Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation

HoneyComb

Findings of EMNLP2024

A Flexible LLM-Based Agent System for Materials Science

11

Foundational evolving-agent framework reimplementing the SciMaster line including ML-Master, X-Master and Browse-Master.

220

Skills-only operating layer for LabOS, with no engine of its own.

1.0k

Foundation Models 14

ChemDFM

Model2024

Developing ChemDFM as a large language foundation model for chemistry

nach0

Model2023

Multimodal Natural and Chemical Languages Foundation Model

DPLM

Model2024

Diffusion Language Models Are Versatile Protein Learners

Efficient 4B genome foundation model trained on 600 Gbp of metagenomic assemblies for de novo sequence generation.

151

scGPT

Nature Methods2024

toward building a foundation model for single-cell multi-omics using generative AI

1.6k

Datasets 55

A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

219

Benchmark and data for scientific question answering over scholarly literature.

163

Structured scientific hypothesis data released with Sparks of Science.

Large open scientific paper corpus for scholarly NLP and retrieval.

Curated bioactivity database for drug-discovery research.

Open chemical information and compound database.

Protein fitness prediction datasets and evaluation resources.

Cleaned 40M-document S2ORC derivative built specifically for language-model pretraining on scientific full text.

All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation Network

302

Official requester-pays S3 bulk access to full-text sources and metadata for the entire arXiv corpus.

CORE

Scientific Data2023

A Global Aggregation Service for Open Access Papers

JSON endpoints for preprint metadata, full-text links and published-version tracking across both preprint servers.

Programmatic access to submissions, reviews, rebuttals and decisions across major machine-learning venues.

QM9

Scientific Data2014

Quantum chemistry structures and properties of 134 kilo molecules

SPICE

Scientific Data2023

SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials

203

Over one million DFT-computed thermodynamic and structural properties of inorganic compounds with API access.

Automatic-flow repository of millions of computed materials entries exposed through a REST interface.

FAIR repository and analysis platform for raw and processed computational materials data across simulation codes.

NIST-hosted integration of DFT, machine-learning and experimental materials datasets plus public leaderboards.

Open platform for reproducible computational materials science with archived AiiDA provenance graphs.

Reference Raman, X-ray diffraction and chemistry data for well-characterized mineral specimens.

MolSSI public quantum chemistry results archive with a Python client for constructing large QC datasets.

Curated protein sequence and functional annotation knowledgebase exposed through a documented REST API.

Experimental three-dimensional biomolecular structures with programmatic search and data web APIs.

Public functional genomics repository of curated expression series, platforms and individual samples.

Uniformly processed functional genomics assays across human and mouse, served through a REST API.

Archive of raw high-throughput sequencing reads across organisms and study types.

Standardized single-cell corpus of tens of millions of cells with an API and in-browser exploration.

Proteomics identifications and mass spectrometry raw data repository hosted at EMBL-EBI.

Machine-readable wet-lab protocol repository, directly useful as a grounding source for protocol-planning agents.

Petabytes of LHC collision and simulated data released with virtual machines and runnable analysis examples.

Gravitational-wave strain data and event catalogs from the LIGO, Virgo and KAGRA observatories.

Space Telescope Science Institute multi-mission archive for Hubble, JWST, TESS and Kepler with programmatic access.

Astrophysics literature and citation database with a full API, the standard discovery layer for astronomy.

Sloan Digital Sky Survey imaging and spectra queryable directly through SQL.

Federated distribution network for CMIP climate model output and related earth-system simulation data.

Earth-observation data portal spanning NASA distributed active archive centers and their APIs.

CERN-operated DOI-issuing general repository hosting a large share of long-tail scientific artifacts and code snapshots.

Open Science Framework project registry and repository covering preregistrations, materials and data.

Large general-purpose research data repository issuing DOIs, widely used across social and life sciences.

Primary social-science data archive holding curated survey, administrative and longitudinal study collections.

German social-science data archive and infrastructure for survey and computational social science data.

Nothing matches those filters.