Distill Sam's judgment.
Forge your own AI model.

KeenForge is an AI agent that distills our CTO's 7-year vision AI workflow — so anyone with trained eyes can train their own expert-grade model. No code. No cloud. You own it — and you earn from it.

CVPR 2026 Demo Nature Portfolio MIT Open Source 100% Local
Get in Touch

What We Do

GPT-4V scores 22% on industrial defect detection. The factory needs 98%+. Generic models hit a structural ceiling in specialized domains. The people who could train expert models — QC engineers, radiologists, agronomists — are locked out by the code barrier. KeenForge replaces the entire ML pipeline with one AI agent. Expert labels and reviews. AI handles everything else — from data cleaning to deployment.


The Theory Behind KeenForge

SAI: AI Must Embrace Specialization via Superhuman Adaptable Intelligence — Goldfeder, Wyder, LeCun & Shwartz-Ziv (2025) — argues the future of AI is not one general model, but thousands of specialized superhuman systems. This is the thesis KeenForge is built on.


The Duo

Lu Gan

Lu Gan

CEO · Domain Expert Turned AI Builder

Nature Portfolio first author. Self-taught from clinical physiotherapy — lived the problem: "I had the domain expertise but couldn't train a model." Now building the tool that eliminates that barrier for everyone.

Sam Li

Sam Li

CTO · 7-Year Vision AI Engineer

Full-stack vision pipeline: data cleaning → model selection → hyperparameter tuning → ONNX deployment. The AI agent we're building is a distillation of his entire workflow.


KeenForge

Not a labeling tool. An AI vision engineer agent. Desktop app. Expert opens it → imports images → draws boxes → AI trains in the background, in real time. The full pipeline runs locally. Your data never leaves your machine.

Import
Auto-Clean
Expert Labels
30-50 imgs
Real-Time
Training
Edge Cases
VLM
Synthetic
Data
Auto-Tune
LLM
Deploy

Green = AI automated  |  White = expert involved

<1s
Label → Train Trigger
17
Model Architectures
100%
Local · Data Never Leaves
MIT
Enterprise License
Existing Tools
Label batch → manually train → wait → load.
Hours to days per loop.
KeenForge
Labeling → background auto-training → sub-second model update.
Producer-Consumer threading.

This is our early-stage prototype — the foundational HITL real-time training framework demonstrated at CVPR 2026. We are now integrating more AI capabilities into the pipeline: VLM edge-case discovery, synthetic data generation, and LLM-driven auto-tuning.

CVPR 2026 Demo GitHub

Publication

Nature · Scientific Data · 2025

"A Curated and Re-annotated Peripheral Blood Cell Dataset Integrating Four Public Resources"

First Author — Curated the dataset, led annotation strategy, and drove the research from clinical insight to publication.