investors about droyd
SAT OCT 10
On this page

Droyd: An Introduction

Open source AI is in vogue. Jensen is collecting signatures. China is topping leaderboards. Americans are open sourcing models. It's happening.

But at the same time it isn't.

Token share of open models is increasing but overall global token spend remains a fraction of overall usage. The vast majority of AI usage happens over products: Codex, Claude Code, and even Grok.

The key word in the last sentence was products.

Products are the last layer of polish that really drive the user experience and why $5 per million tokens is palatable and why you keep coming back for more.

That polish comes from a highly optimized feedback loop owned by the frontier labs across the model, harness, and product. The frontier models have gotten so optimized for their harnesses that they perform materially worse on benchmarks outside of their intended harness.

Open source AI completely lacks the optimization feedback loop between model and harness.

Operating relatively independently, Chinese labs train the models. American providers host them. And one of the 5 major open source harnesses deliver them to users.

Each player in the open source AI stack is playing their own game.

What's missing, and what drives the whole flywheel around, is data. Data in the form of task examples or, specifically, trajectories which teach models to natively drive a variety of harnesses in complex, long-horizon environments.

Open training data geared for open harnesses will dramatically improve model performance inside the very tools users are using to interact with AI. The challenge is generating good data across all the top harnesses.

Droyd is completing the open source AI flywheel - we are building an incentivized data factory using competitions to generate continuous trajectory data across top benchmarks and harnesses.

Droyd: The Eval Competition Platform

Droyd is a platform for hosting open eval competitions. For each eval benchmark competition we host, users submit optimized agent prompts, tools, and skills in hopes to out score other submissions on the evaluation tasks. Higher scoring submissions win a proportional share of the prize pool.

We target hosting a variety of benchmark competitions across domains such as coding, finance, commerce, biology, legal, and knowledge work with a focus on complex long-horizon tasks.

As users iterate to find good submissions, each test run produces trajectory data for the particular model and harness solving the task. The better the submission, the better the agent is able to utilize the harness and specialized knowledge to solve a given task producing better training data.

Over time, a collection of trajectory data across domains, harnesses, and models is assembled, scored, and sorted for downstream training.

Competitions & Rewards

Each competition on Droyd is an eval competition in which new races are conducted every day.

Users iterate locally with their coding agent of choice seeking to maximize scores on test datasets and submit their best configurations. Submissions carry an entry fee - typically $2-$5 - which goes to fund the prize pool. Once the race submissions are scored, the top submissions are paid out along the defined payout curve.

Open source data pipeline from local optimization through competition rewards and open source submissions

The user experience of competing in a competition can be envisioned as a pipeline:

  • Local Optimization: using the Droyd CLI, users conduct an auto-research loop with a variety of test training sets to achieve a worthwhile score. Evaluations are hosted on Droyd thus minimizing the hardware requirements of users.
  • Submission: Using the best local optimized package of prompts, skills, and tools, users pay the submission fee to the competition's Solana race contract. Payment happens with the user's Droyd wallet via API request and goes to fund the prize pool.
  • Evaluation Race: Races close at a set time at which all the submissions are queued for evaluation. Each result is signed and saved for open download and inspection.
  • Payout: Once all submissions are scored at the end of a race, a signed manifest of the top submissions is published to the race contract and winners can claim their earnings. Payout curves are typically designed to be a power-law weighted curve paying out the top 20% of submissions.
  • Open Source Release: winning submissions are released after races to promote rapid, global optimization and transparency.

The pipeline is effectively a global auto-research loop where local agents optimize within the day and check-in their best improvements to the global repository. During both the local optimization and evaluation race, the submission runs in Droyd environments and scored trajectory data is captured.

Competitions are expected to range across disciplines. Coding, finance, legal and more. A roadmap of competitions is summarized as:

  • MazeBench: long-horizon planning and strategic optimization inside of a complex puzzle environment
  • SWEBench (variants): advanced coding tasks to capture ability to plan, use subagents, and execute long-horizon tasks
  • LegalBench: combination of top legal benchmarks capturing agentic abilities on complex legal tasks
  • ShoppingBench: commerce tasks designed to measure an agent's taste and tool use in consumer shopping situations.

The core muscle in hosting a good competition is task generation to combat overfitting and promote higher degrees of diversity in the resultant datasets.

Environments

Droyd offers hosted tasks execution as we have prebuilt environments for each dataset and harness making local iterations and data generation as simple as calling an API.

Hosted evaluation pipeline turning local optimizations into ranked agent traces and data products

Each task within a competition can have a different environment with its own code dependencies, compute requirements, GPU requirements, etc. and to manage this all locally as a user is far too complex. Droyd pre-configures each task with an environment and harness image which is spun up in a sandbox upon evaluation.

This results in both lower barriers to entry for users and a centralized platform to capture and format trajectory data for downstream consumption.

Users receive a $10 compute credit at sign up and from then on, purchase credits to fund their compute balance which is charged per second of runtime.

Competition Tooling

Droyd offers a CLI toolkit and plugin for local coding agents such as Codex, Claude Code, Hermes and more. The CLI enables coding agents to quickly understand the competition rules and provides the auto-research scaffolding needed to build an optimized submission. The goal is to lower the barriers to entry so that more users are able to participate and the resulting datasets are more diverse.

One core aspect in helping local coding agents navigate long-horizon optimizations is the Droyd auto-research API.

Droyd Experiment API interface for recording hypotheses, reasoning traces, and experiment results

Local coding agents automatically have each of their experiments recorded with thoughts, hypotheses, and reflections so that it can remain coherent over many experiments. Each experiment is auto annotated with the evaluation results and provides agents a way to query for what's working and see why without having to rebuild its own auto-research framework.

Trajectory Store

Completed evaluation runs produce trajectories which go into the Droyd trajectory store for purchase. Each trajectory is labeled and scored so that buyers are able to segment data purchases by harness, task difficulty, score percentile and more.

Trajectories will be available for purchase via onchain payments via x402 and MPP which will go directly to the Droyd treasury.

We see trajectory data customers as

  • AI Labs: frontier trajectory data across variety of harnesses and domains serves the need for SFT warm ups in model post-training
  • Harness and API Developers: harness developers can filter the dataset for tasks using their apps to select tasks where agents perform both well and poorly to improve their products.

By generating the trajectory data via user-funded competitions, the trajectory data can be sold to downstream consumers for fractions of the cost to generate the data.

Revenue Model

Droyd earns revenue from three core areas

  • Trajectory Downloads: per trajectory with variable costs for quantity and top performing, recent data
  • Competition Platform Fee: 10% of the prize pool for any given competition is sent to the treasury
  • Evaluation Compute: each evaluation run is billed per second of compute relative to the size of compute and GPUs consumed

The revenue strategy is to maximize trajectory download revenue as it has minimal marginal cost and has the highest scalability while using the other revenue lines to cover costs.

Droyd: The End State

Continuous learning arises when novel, frontier-advancing tasks can be generated faster than the training pipeline can consume them.

But simply asking an agent to conjure a new task isn't so simple. If the agent can one-shot a new, frontier-advancing task, why can't it also solve that task with the current model?

To successfully have an agentic task generation pipeline (and thus continuous learning pipeline), a few ingredients are required:

  • Base Data Generation: many tasks require environments with various seed data in simulated environments (think simulated slack company) and a prerequisite is to have engines to simulate this data and interfaces for tasks.
  • Solution Proofs: tasks need to be solvable. Either tasks follow a deterministically measured tasks or judge-based scoring in which the target solution should be known.
  • Success Rate Testing: tasks are only valuable if a tiny fraction of the agents solving them can accomplish it. Too easy, it doesn't offer anything for training. Too hard, there are no success cases to train against. The solution here is to be able to A/B test tasks in real environments - such as Droyd competitions - to know which should be discarded, modified, or promoted.

With Droyd, we envision a flywheel where the more users competing in competitions, the more scale we can A/B task generations against, the greater scale we can sell data, and thus the more rewards we can offer to attract even more users.

Droyd flywheel connecting more agents, trace data, higher scores, and rewards into a continuous reinforcement learning pipeline

At its end-state, Droyd becomes a critical piece in the continuous training pipeline constantly testing and developing new tasks to train frontier models against.