TradeForge GPU compute, learning loop and swappable models (designed, not built)

The idea and the design

click to collapse or expand
Designed and owner-approved (decisions G1 to G15, revisions R1 and R2, 2026-10-07). Nothing is built yet.

The idea

TradeForge tests trading strategies by running huge numbers of calculations: simulations of thousands of possible futures, replays of years of market history across thousands of stocks and settings, model training and portfolio sizing. Most of that is the same calculation repeated on different inputs, which is exactly what graphics chips (GPUs) and other specialised chips do fast.

The idea is to make speed a property of the platform, not something each feature builds for itself. TradeForge describes the work it needs. The platform decides which machine runs it and which kind of chip does the math. AI models get the same treatment: each lives in its own container, and a setting decides which model answers which kind of question.

The design also draws the learning loop that already exists. Each night, AI models propose new strategy ideas, ordinary code tests them, and every result is kept. The next night's ideas are written with all past results in view, so the system gets better at choosing what to try. Faster chips make each test cheaper.

The forge ecosystem in plain words

New to the forges? These are the names used on this page.

The forge ecosystem
A family of separate code projects, each called a forge, that each own one job. Apps such as TradeForge reuse them instead of rebuilding the same plumbing.
TradeForge
The trading app this design is for. It researches strategies, checks them, and places trades.
forge-native
Holds fast low-level code that many forges share. In this design it gains the compute layer: the parts that decide where a job runs and which chip does the math.
forge-core
The common base every forge builds on. In this design it owns the single standard format for talking to AI models.
forge-gateway
A switchboard for AI models. A forge asks for a job title such as 'build' or 'audit', and the gateway decides which model answers.
configforge
Holds settings per environment, such as 'on the laptop, use these machines'.
devopsforge and forge-deploy
Generate and store the deployment instructions that put programs onto Kubernetes.
llmforge
Decides when an AI model should be trained or slimmed down, and runs that training.
backboneforge
Holds the shared knowledge graph: facts and documents, linked to each other.
forge-ingest
Brings documents, PDFs and web pages into the knowledge graph.
agentforge
Runs AI agents and GraphRAG, which looks things up in the knowledge graph to answer questions.
forge-console
The web console for the whole ecosystem, including the queue where people approve new knowledge.
Kubernetes
Software that runs programs in containers across one or many machines. A Kubernetes Job is a one-off task that runs to completion.
Container
A packaged program plus everything it needs to run, so it behaves the same on any machine.

The design at a glance

Two switch points
One decides where a job runs: the local cluster on the Mac, a GPU box you might buy, or rented cloud GPUs. The other decides which chip does the math: ordinary CPU, Apple, NVIDIA, AMD, Intel or others. Code that wants work done never names a machine or a chip.
Kubernetes everywhere
Every run, even on the laptop, is a Kubernetes Job. Local and cloud work the same way.
Data by reference
Jobs carry labels pointing at data, never the data itself. Each machine fetches its own copy.
Trustworthy results
Fast-chip code must match the plain CPU answer before use. A certified result records the exact chip and software it ran on, so it can be replayed exactly.
The DAG split
The live daily scan stays inside the app because it is on the trading path. All batch scan work, such as what-if sweeps and re-checks over history, runs as jobs.
The learning loop
AI models propose, code tests, every result is kept, an AI critic reviews, humans promote. Models improve by reading past results. Training a small model on those results is a later phase.
Knowledge
Documents flow into a knowledge graph. The idea-proposing model looks things up there before proposing.
Swappable models
One standard request format, one translator per vendor API, and a switchboard that maps job titles to models. Changing a model is a settings change.

Step by step: how work and data flow

Flow 1: a compute job

  1. TradeForge parts, top left. A part of TradeForge, say Kelly Monte Carlo, has heavy math to do. It does not pick a computer or a chip. It fills out a request instead.
  2. Compute job. That request is the compute job. It says what kind of work is needed, such as 'batch array math' or 'solve an optimisation problem'. It gives a rough size. It lists which data to read and where to put the results, as data references rather than the data itself.
  3. configforge and Router. The job goes to the Router. The Router asks configforge for the list of machines this environment may use, in preferred order, and picks the first that can do this kind of work.
  4. Capability registry and Finance kernels. The registry is a table of which code can do which kind of work. TradeForge's own math, such as the Monte Carlo simulator, has registered itself there. A plain CPU version always exists, so every job can run somewhere.
  5. ComputeTarget, the first switch point. This decides where the job runs: the local cluster on the Mac, a GPU box you might buy later, or a rented cloud GPU pool that switches off when idle.
  6. Kubernetes Job and Compute images. Whichever machine is chosen, the job is packaged as a Kubernetes Job and runs inside a container image built for that machine's chip.
  7. Accelerator backend and Chip plug-ins, the second switch point. Once the machine is chosen, this decides which chip does the math: CPU, Apple, NVIDIA, AMD, Intel or another family. It loads the matching version of the code.
  8. Certification, the red boxes. Before fast-chip code may be used, Admission checks that its answers match the CPU version closely enough. If the result will be certified for trading, the CertificationReport records the exact chip and software, so the run can be replayed exactly later.
  9. Data plane. While the job runs it needs data. Data references name each piece of data by store, location and version. The per-target resolver turns a reference into a real local copy. In the cloud it pulls files from object storage once and keeps them on the machine's fast disk.
  10. Results come back. The job writes its results as new data references, which return to the caller the same way the inputs arrived.

Flow 2: the learning loop

  1. Coordinator. Each night the Coordinator, which is plain code and not an AI, starts a cycle. It also runs safety switches that stop the loop if something looks wrong.
  2. Proposer. An AI model working under the 'build' job title reads past results and looks up relevant knowledge through GraphRAG. It then proposes up to 32 strategy variants to test.
  3. Backtest evaluation. Ordinary code tests every variant against market history. This is the heaviest work in the loop, so it is sent out as compute jobs and follows Flow 1.
  4. Verdict registry. Every result is recorded permanently, including failures. Keeping failures stops the system from fooling itself by forgetting what did not work.
  5. The loop closes. The next night, the Proposer reads the Verdict registry again. That is how the system learns: its ideas improve because it sees everything tried so far. The AI model's own weights do not change.
  6. Critic. A stronger AI model working under the 'audit' job title reviews each cycle and flags results that look too good to be true.
  7. Coordinator and Human review. Code updates the scores and a promotion ladder: research, then candidate, then certified. Only humans can promote a strategy to paper or real trading.
  8. Later phase: llmforge. Planned but not built: llmforge trains a small proposer model on the Verdict registry, as a compute job. The trained model is registered like any other and still faces the same fixed tests.

Flow 3: asking an AI model

  1. The callers. Parts of TradeForge that need an AI model, such as the Proposer, the Critic, the Forecasters and the chat, ask by job title, such as 'build' or 'audit'. They never name a specific model.
  2. forge-gateway. The gateway looks the job title up in its registry entries. Each entry says where a model lives, which API language it speaks, and the model's name. A job title lists several models in order, so if one is down the next is tried.
  3. forge-core contract. The question is written in one standard format, so no vendor's quirks leak into TradeForge.
  4. The translators. Each translator turns the standard format into one vendor's API language: OpenAI, Anthropic, Ollama, Bedrock or Gemini, or the Open Inference Protocol for number-predicting models. A new kind of API means one more translator and nothing else.
  5. Endpoints and model servers. The request goes to the model's address, using its stored credential. The model answers from NVIDIA's model containers on a GPU machine, from Ollama on the Mac, from a paid hosted service, or from an embedding model used by GraphRAG. The answer travels back the same way.

Flow 4: knowledge

  1. forge-ingest. Documents, PDFs and web research come in. People approve new items in forge-console's review queue.
  2. Knowledge graph. backboneforge stores the knowledge as linked facts.
  3. GraphRAG retrieval. agentforge searches the graph and gathers relevant passages. It uses embedding models, reached through the gateway like any other model, to find related text.
  4. Into the loop. The Proposer recalls knowledge through GraphRAG before proposing. A link that would start a learning cycle whenever new knowledge arrives is planned but switched off for now.

Trade-offs

ChoiceWhat it buysWhat it costs
Kubernetes for every run, even localOne way of running things everywhere, so a job that works on the laptop works in the cloud.Every local run pays container start-up time. And the Mac's GPU is out of reach from containers under Docker Desktop, so the local cluster does its math on CPU only.
Open to any chip familyNew chips can be added later without changing callers.Every chip family needs its own version of each calculation, its own container image and its own admission tests. Only CPU and NVIDIA are likely to be built at first.
Two switch points instead of oneCallers never change when hardware changes.More moving parts, and two or three versions of each calculation to keep in agreement.
A CPU answer for everythingEvery fast-chip result can be checked, and work still runs when no GPU is available.No calculation can be GPU-only. GPU libraries such as cuOpt need a separate CPU equivalent whose answers differ in their own ways.
Certification pinned to chip and softwareReplays stay exact, so 'pin everything' still holds.When a cloud provider retires a GPU type or updates drivers, old certifications must be re-run.
Live scan stays in the appThe trading path keeps no extra delay and no new way to fail.The live scan does not get faster. If it grows heavy, it will need its own design.
Data by referenceJobs stay small, and every result names exactly which data it used.The referenced data must already be in the cloud. The data lake is about 464 GB, roughly 11 hours to upload at 100 Mbps, and where it lives is still undecided.
One owned model formatNo vendor's API shape leaks in, and a new API is one translator.The forge team maintains every translator. Vendors add features faster than one person can map them.
Learning through context firstWorks today, needs no training, and cannot drift away from the fixed tests.Context has a size limit. As the Verdict registry grows, the Proposer sees a summary rather than everything.

Possible flaws

  • The speed-ups are not measuredNo calculation has been timed on a GPU yet. The decade walkers branch a lot and move day by day, which GPUs handle poorly. Reading the data may take longer than the math. Each phase should start by timing the CPU version and one GPU prototype.
  • Faster tests do not mean more discoveriesThe loop counts every variant ever tried and raises the bar as the count grows, so lucky results are not mistaken for skill. Cheaper backtests invite more variants per night, which raises that bar faster. Speed must buy better-chosen variants, not more of them.
  • A model trained on its own results can narrow its searchIn the later phase, a proposer trained on past verdicts may keep proposing variations of past winners. The loop already requires diverse batches of variants, and that rule must stay in force.
  • Phase 1 may show no gainAt 10,000 paths, the Kelly simulation may run as fast on a CPU with good array code. It still proves the plumbing works, but not that GPUs pay off.
  • Cheap chips are weak at high-precision mathThe Apple GPU has no 64-bit floating point. Low-cost NVIDIA cards run it at a small fraction of their normal speed. If results must match the CPU at 64-bit precision, cheap chips may never pass admission.
  • Ledger positions assume nothing is deletedA data reference such as 'trade events up to position N' only stays stable if earlier rows are never deleted. Past clean-ups did delete rows, which would silently change what old references mean.
  • Most chip plug-ins are hypotheticalA switch point is only proven when at least two real options sit behind it. Today there is no GPU box, no cloud GPU pool, and no AMD or Intel code. Build each only when the hardware exists.
  • The Mac's local models conflict with Kubernetes everywhereLocal AI models run on the Mac itself so they can use its GPU. Moving them into the cluster would make them CPU-only. This is still open.
  • Routing ignores cost and delayThe Router picks the first capable machine in the list, so a tiny job can go to a cold cloud GPU. The size field in a job is not defined yet, so it cannot help.
  • Licensing is unverifiedRunning NVIDIA's model containers yourself in production may need a paid NVIDIA licence.

Click any box in the diagram below for its own explanation, or use About this design at the bottom left.

TradeForge GPU compute, learning loop and swappable models (designed, not built) An architecture diagram generated by Archify. Forecasters · numeric models (Chronos) · TradeForge (the trading app) Forecasters numeric models (Chronos) Chat + summarizer · asks by role · TradeForge (the trading app) Chat + summarizer asks by role Registry entries · endpoint + dialect + model · Model access Registry entries endpoint + dialect + model forge-gateway · role → model list · Model access forge-gateway role → model list forge-core contract · one standard format · Model access forge-core contract one standard format OpenAI chat · Model access OpenAI chat Anthropic messages · Model access Anthropic messages Ollama native · Model access Ollama native Bedrock / Gemini · Model access Bedrock / Gemini Open Inference · numeric models · Model access Open Inference numeric models Endpoints · address + credential · Model access Endpoints address + credential NIM · NVIDIA GPU · Model servers (one model each) NIM NVIDIA GPU Ollama / llama.cpp · Mac host · Model servers (one model each) Ollama / llama.cpp Mac host Hosted APIs · paid per call · Model servers (one model each) Hosted APIs paid per call Embedding models · for GraphRAG · Model servers (one model each) Embedding models for GraphRAG Knowledge graph · backboneforge · Knowledge Knowledge graph backboneforge GraphRAG retrieval · agentforge · Knowledge GraphRAG retrieval agentforge forge-ingest · documents + web in · Knowledge forge-ingest documents + web in Human review · only humans promote · Self-learning loop (nightly) Human review only humans promote Coordinator · cycles, ladder, breakers · Self-learning loop (nightly) Coordinator cycles, ladder, breakers Proposer · build-role model · Self-learning loop (nightly) Proposer build-role model Backtest evaluation · code, ≤32 variants · Self-learning loop (nightly) Backtest evaluation code, ≤32 variants Critic · audit-role model · Self-learning loop (nightly) Critic audit-role model Verdict registry · every result, kept forever · Self-learning loop (nightly) Verdict registry every result, kept forever Trained model server · + registry entry · Model training Trained model server + registry entry llmforge · decides when to train · Model training llmforge decides when to train Kelly Monte Carlo · phase 1 · bet sizing · TradeForge (the trading app) Kelly Monte Carlo phase 1 · bet sizing Decade walkers · phase 2 · batched backtests · TradeForge (the trading app) Decade walkers phase 2 · batched backtests Portfolio optimiser · phase 3 · cuOpt + cuML · TradeForge (the trading app) Portfolio optimiser phase 3 · cuOpt + cuML DAG batch work · what-if, shadows, re-cert · TradeForge (the trading app) DAG batch work what-if, shadows, re-cert Compute job · what, size, data refs · forge-native compute layer Compute job what, size, data refs configforge · targets per environment · forge-native compute layer configforge targets per environment Router · first capable target · forge-native compute layer Router first capable target Capability registry · work type → code · forge-native compute layer Capability registry work type → code Finance kernels · TradeForge's math · forge-native compute layer Finance kernels TradeForge's math ComputeTarget · seam: WHERE · forge-native compute layer ComputeTarget seam: WHERE Compute images · one per chip family · Kubernetes everywhere Compute images one per chip family Kubernetes Job · every run, every target · Kubernetes everywhere Kubernetes Job every run, every target Local cluster · kind on Mac · CPU only · Kubernetes everywhere Local cluster kind on Mac · CPU only Local GPU box · k3s · future · Kubernetes everywhere Local GPU box k3s · future Cloud GPU pool · scales to zero · Kubernetes everywhere Cloud GPU pool scales to zero Accelerator backend · seam: HOW · any chip · forge-native compute layer Accelerator backend seam: HOW · any chip Admission · must match CPU · Certification Admission must match CPU CertificationReport · pins chip + versions · Certification CertificationReport pins chip + versions Chip plug-ins · CPU, Apple, NVIDIA, AMD, Intel, TPU… · forge-native compute layer Chip plug-ins CPU, Apple, NVIDIA, AMD, Intel, TPU… Data references · store + path + version · Data plane Data references store + path + version Per-target resolver · finds a local copy · Data plane Per-target resolver finds a local copy Object storage · cloud copy · Data plane Object storage cloud copy Node NVMe cache · cached by hash · Data plane Node NVMe cache cached by hash Live scan (DAG) · stays in API pod · TradeForge (the trading app) Live scan (DAG) stays in API pod Postgres ledgers · version = log position · Data plane Postgres ledgers version = log position Parquet lake · version = content hash · Data plane Parquet lake version = content hash DuckDB files · version = content hash · Data plane DuckDB files version = content hash Snapshot export · other stores → Parquet · Data plane Snapshot export other stores → Parquet batch DAG llm_finetune backtests route target order lookup selects register submits runs on runs on runs on image resolves plug-in admit kernel pass refs in + out resolve cloud pull cache predictions start cycle ≤32 variants verdict past results review findings ladder + digest role: build role: audit recall knowledge add graph embeddings + answers new knowledge (deferred) role role lookup standard format becomes training data (later) TradeForge (the trading app) Self-learning loop (nightly) Knowledge forge-native compute layer Certification Kubernetes everywhere Data plane Model access Model servers (one model each) Model training Legend Caller or person Component or job Store, reference or config Runtime: cluster, image, model server Certification Seam or contract Dialect adapter