JK/ LOS ANGELES
PT --:--:--
SPIKING NEURAL NETWORK · LIVE · MOVE YOUR POINTER

JoshuaKezzer

AI/ML ENGINEER· SYSTEMS ARCHITECT· VFX ARTIST

“if it’s slow, profile it.
if it’s magic, disassemble it.”

SPIKES/S0
NEURONS0
SYNAPSES0

Position

EST. LOS ANGELES

I build AI systems, trading bots, and automation tools — and I run them on my own hardware.

I started in VFX and spent about ten years there: compositing, lighting, and the tools that keep a render farm moving. Then three years on trading systems, where the software handles real money and a wrong assumption doesn’t crash — it just quietly loses some of it. Now I build AI systems: agents that run without supervision, models served on machines I own, and the plumbing underneath both.

What carries across all three is that I don’t treat the layer below mine as somebody else’s problem. When a model is slow, the fix is usually in how it stores context while it generates, not in buying a bigger model. When a DeFi protocol advertises a big yield, the real number is what’s left after transaction fees and after you price in the odds the contract gets drained. “It’s slow” is a claim you can check: run a profiler, read the trace, find the actual bottleneck. Most systems I open are running 2–4× slower than they need to because nobody did that.

I write specs too. Fifty-page architecture documents a team can build from — a recommendation system for a streaming service, the data model that has to sit underneath it, and the privacy paperwork that has to hold up in front of a lawyer. Same skill set as building the thing, so I do both.

The network up top isn’t decoration. It’s a live simulation of neurons: each one builds up charge until it crosses a threshold, fires, and passes the signal down its connections with a real travel delay. The strip below plots every spike as it happens. Move your pointer across it and you’re poking the network with an electrode.

Arc

3 ERAS · 2011–NOW
2011 – 2021·COMPOSITING, LIGHTING, TOOLS

VFX & Pipeline

I worked as a compositor and lighting artist first, then moved into building the tools around the work. The bottleneck was never the artist — it was a pipeline that made you wait forty minutes to see one frame. So I wrote the batch job system, the plumbing that made renders come out identical every time, and the monitors that caught a dead render node before a producer noticed. Ten years of that taught me to always look at what’s happening underneath the tool I’ve been handed.

CompositingLighting & Look-Dev Pipeline ToolingRender Orchestration Color
2021 – 2024·TRADING SYSTEMS

Crypto & Fintech

Prediction markets, DeFi, and retail arbitrage — all of it running live with real money on it. PolyBrain grew to about 66,000 lines of Python and Rust: a group of language models estimating what a Polymarket contract is actually worth, eight services reading the order book, and a Rust component that got the scan loop from roughly a second down to 5 microseconds once it was clear Python couldn’t keep up. YieldBrain scores six DeFi protocols on a seven-part risk model and decides how much money goes into each one, weighted toward the chance of losing everything rather than the average case. Finding the edge was the fun part. Building the safety limits that shut it all down when something goes wrong took longer.

Prediction MarketsDeFi Yield Arbitrage EnginesSolidity RustRisk & Circuit Breakers
2024 – PRESENT·AGENTS, MODELS, INFRASTRUCTURE

AI/ML Systems

HQ runs 31 tools across four machines linked by a private VPN — a 9950X/RTX 5090 workstation, an i9-14900K/RTX 4090 second box, a Raspberry Pi 5, and a laptop. All of it local, because renting inference means you never quite know what the machine is doing when it’s busy. NeuroBrain predicts how a person’s brain responds to video and audio, then draws the result onto a 3D model of the cortex: 108,219 points per frame from one pipeline, or a combined video/audio/text model that produces 1,000 brain regions for every 1.49 seconds of footage. Alongside that I write architecture specs on contract — a recommendation system for a streaming platform covering film, music, podcasts, and live events, tested against a simulated population of fake users generated by a 235B-parameter model, so the ranking could be evaluated before any real users existed.

Multi-Agent OrchestrationLocal LLM Serving Recommender SystemsWireGuard Mesh ComfyUI PipelinesSelf-Hosted PaaS

Focus

6 DOMAINS
MARKETS & CAPITAL

Trading Systems

Bots that trade prediction markets and move capital between DeFi protocols, with hard limits that stop them when losses hit a number I set in advance.

Prediction MarketsDeFi Yield Rotation Arbitrage ScoringCircuit Breakers
MODEL SERVING

LLM Infrastructure

Getting language models to run fast on hardware you already own — serving with vLLM, shrinking models so they fit in less memory, and tuning how they hold context while generating.

vLLMKV-Cache QuantizationLocal-First Inference
RETRIEVAL & RANKING

Recommender Systems

Making movie, music, and podcast recommendations work out of one system — including for items that are brand new and have no viewing history behind them yet.

Two-Tower RetrievalCross-Domain Matching Ranking ModelsCold Start
ORCHESTRATION

Autonomous Agents

Multi-agent orchestration that scores sessions across seven quality signals, detects file-edit conflicts, and ships REST, SSE, and Prometheus metrics with a CI gate.

Multi-Agent Systems31-Tool HQ Session ScoringConflict Detection
NEUROSCIENCE

Brain Mapping

Predicting how people respond to media by mapping neural activation across 108K cortical surface points — fusing vision, audio, and language models against a 47-region atlas.

fMRI PredictionMulti-Modal Fusion Cortical MappingPyTorch
NETWORKING

Infrastructure

A four-machine encrypted mesh with self-healing watchdogs, GPU-distributed job scheduling, and a fleet orchestrator that scores every agent session across seven quality signals.

WireGuardFleet Orchestration GPU SchedulingSelf-Healing Mesh

Stack

35 TOOLS / 5 LAYERS

Languages

PythonRustC++ TypeScriptSolidity CUDA CPowerShell

Models & Inference

PyTorchvLLMOllama llama.cppTensorRTHuggingFace sentence-transformersDiffusersComfyUI

Data & Services

PostgreSQLClickHouseRedis SQLiteKafka / RedpandaFastAPI PlaywrightGradio

Markets & Chain

web3.js / ethersFoundry Polymarket APIDeFi Protocols

Infrastructure

DockerLinuxWireGuard CoolifyCloudflare Tunnel NginxAWS

Work

LOCKED · KEY REQUIRED
18 SYSTEMS · 6 MODELS · SEALED

RESTRICTED

ENTER ACCESS KEY

Project write-ups are private. Ask Josh for the key.