Three years of inference engines, GPU kernels and production services for a distributed AI compute network. Before that, the CUDA backend of a Python-to-GPU compiler. Every live demo under work runs client-side, in this tab. No backend. No API key.
Teams size GPUs off the model file and then OOM in production, because the KV cache — not the weights — is what scales with context and concurrency. This computes the real total for actual architectures and tells you which cards it fits on.
Tensor decompositions get published as math and almost never as something you can run. This is a browser-oriented approximate extension of my generalized ℒ-product paper: overlapping spatial tiles, exact per-tile SVD, fanned across a worker pool.
Motion invisible to the eye is already in the video; it only needs amplifying at the right temporal band. Eulerian magnification falls out of the tensor codec as a reconstruction-time knob, so the same decomposition that compresses also reveals.
Every coding-agent demo is a thin client for someone else's API key. This one runs a 1.5B model on your own GPU over WebGPU, edits a virtual filesystem, and runs the tests for real in a Pyodide sandbox — nothing leaves the tab.
Five local quantized models — one large, four small — on a curriculum with checkable answers. No fine-tuning: the only state a small model carries between sessions is a notes file it writes itself, and an auto-graded quiz each session says whether that moves the score.
A. El Hachimi, M. Elalj, K. Jbilou, A. Ratnani. “Generalized ℒ-Product for High Order Tensors and Applications Using GPU Computations.” Mathematical Modeling with Modern Applications (M3A 2024), Springer PROMS vol. 497, 2025.
doi:10.1007/978-3-031-89041-3_6