Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Built at Berkeley

A living, evidence-gated database of UC Berkeley founders & builders.

Tracks Berkeley-affiliated founders across the major startup pipelines (Y Combinator, a16z Speedrun, The House Fund, Berkeley SkyDeck) alongside a curated roster of notable Berkeley builders — open-source authors, lab researchers, and hackathon winners. Every entry is primary-source verified: a real link confirming the UC Berkeley tie plus a concrete thing the person built.

Built at Berkeley — builders view

↑ Builders view — illustrative sample data; the real roster is private.

Note on data: this public repo is the engine and design, not the roster. The person-level data (state files, founder/builder records, generated dashboards) is kept private and excluded via .gitignore. The live dashboards are not published to a public URL and carry noindex.

Highlights

  • Evidence-gated — nothing is listed without a primary source; the scorer is tested against real false-positive traps and never invents a person or company.
  • Two views — a companies database and a builders roster, each with search, category/source filters, and a compact spreadsheet mode.
  • Standardized bios — founders render as a clean Major · degree · year token (e.g. EECS '25, Biology PhD '22, Haas MBA).
  • Provenance tags — each builder is badged by the lab / accelerator they came through (BAIR, Sky Lab, RISELab, AMPLab, AUTOLab, SkyDeck, …).
  • Self-contained — dependency-free, theme-aware static HTML; works from file:// or any static host.

The two views

  • Companies — every startup with a confirmed UC Berkeley founder, aggregated across six accelerator sources. Founders shown one per line with a standardized bio; filter by source, search by founder, toggle a spreadsheet layout.
  • Builders — a curated roster of notable Berkeley builders, tagged by lab / accelerator and filterable by category, in card or spreadsheet layouts.

How it works

pipeline

An owner-run pipeline — nothing runs server-side:

seed roster → sweep (public-snippet search) → score (UC-Berkeley-only,
with false-positive traps) → forecast → diff → persist → regenerate dashboards

The scorer is the interesting part: it counts UC Berkeley only and is tested against real false-positive traps that must not count — Berkeley College (an unrelated school), Berkeley Lights (a company), a "located in Berkeley, CA" address, and Berkeley Lab / LBNL (the national lab, not the university). Human- verified tiers are never silently downgraded by a weak automated sweep, and any founder the sweep can't attribute to a known company is parked for a human to resolve — the agent never invents a company or a person to fill a gap.

Engine (public)

agent/
  sweep.py     the agent: sweep → score → forecast → diff → persist → redraw
  scoring.py   UC-Berkeley-only scoring with the false-positive traps + self-test
  search.py    pluggable search backend (Brave / SerpAPI / keyless DuckDuckGo)
.github/workflows/watch.yml   scheduled sweep + GitHub Pages publish (opt-in)

Run the scorer's self-test:

python agent/scoring.py     # trap/signal self-test

The seed rosters, aggregated state, curated builder/major records, and the dashboard generator live in modules that are not part of this public repo.

Tech

Pure Python standard library for the pipeline; self-contained, theme-aware, dependency-free static HTML dashboards; optional Supabase backend for community contributions.

License

All rights reserved. Data is private and not licensed for redistribution.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages