A hardware research lab run by AI coding agents

ChipLab designs and verifies AI-accelerator hardware. The engineering is done by AI coding agents (Claude and Codex) and directed by the lab's founder. Every result an agent produces is checked by a different agent before it is accepted.

The lab works only with free and open-source tools: Yosys, Verilator, Icarus Verilog and OpenROAD. Its work so far is design, simulation and open physical-design flows. No chip has been fabricated.

ChipLab is operated by Donmakino LLC (dba ChipLab), a Virginia company.

How it works

  1. Issue

    Each piece of work starts as an issue with one question and a check that can fail.

  2. Agent

    One agent claims the issue and works on its own branch, in a pinned open-source toolchain. It opens a pull request with a short result note: the claim, the numbers, the method and the limits.

  3. Independent review

    A different agent recomputes the claim at the exact commit and posts a verdict that names that commit: confirmed, or a specific discrepancy. Agents never review their own work.

  4. Merge train

    A pull request merges only when its current commit carries a confirmed verdict and its checks pass. A new push needs a new verdict. Failed attempts and negative results stay on record.

The record so far

From the lab's own GitHub history, 23 September to 4 October 2026 (12 days). The first days ran under a temporary self-review policy. Independent review has dominated since 29 September, and the self-review policy formally ended on 3 October.

Pull requests merged
1,227 of 1,365 opened
Independent review verdicts
1,221
Verdicts that found a discrepancy
23.3% 263 of 1,131 decisive verdicts
Reviewed pull requests that needed a fix round
170 of 731
Merged results that report a negative or failed outcome
at least 45 of 446
Merged pull requests with an independent verdict
688 of 1,227 (56%)

What the reviews catch

A hand-labelled sample of 48 of the 263 discrepancy verdicts, by the main blocking finding.

FindingShare
Broken check. A test or validator accepts bad input, never runs, or fails at the reviewed commit.33%
Unsupported claim. The text asserts more than the evidence shows.25%
Wrong number. A value, bound or unit is wrong.17%
Stale reference. A cited commit, file or assumption is out of date or unreadable.15%
Other. Not ready, out of scope, or incomplete evidence.10%

Some of the broken-check findings show the author's own check accepting input it should reject: for example, a verifier that accepted a duplicated data shard as a complete packet. The reviews test the checks, not only rerun them.

Open source

bf16-fp32-ops

A BF16 × BF16 → FP32 multiplier and an FP32 adder in plain, synthesizable SystemVerilog, released under the Apache-2.0 license. The modules match a Python reference model bit for bit on about 884,000 test vectors (604,032 multiplies and 279,600 adds), on both Icarus Verilog and Verilator. All six seeded bugs in its mutation tests are caught, and the code is lint-clean under Verilator -Wall. The tests are not exhaustive and are not a formal proof, and no area, timing or power figures are claimed. The RTL, model and tests were written by the lab's agents.

Contact

[email protected]