Mouad El Alj
Software Engineer · Python, C/C++, Rust, CUDA
Systems engineer working from GPU kernels up to distributed services: C/C++ and Rust for performance-critical code, Python for machine learning and tooling. Three years on inference engines, accelerator code and production services.
Experience
Software Engineer, AI Infrastructure · Hyperspace
Apr 2023 – Feb 2026
Distributed compute network running machine-learning workloads on contributed consumer hardware.
- GPU and accelerator programming. WGSL/WebGPU compute shaders with bit-identical CPU implementations, so GPU work can be verified on CPU. Tuned memory throughput and batch sizing across hardware; maintained CUDA and Metal builds.
- Low-level systems and native integration (C/C++, Rust). Wrote a C API over a large C++ inference library plus safe Rust bindings, embedding it in the host process instead of a separate server. Cross-platform builds for Linux, macOS and Windows, including GPU toolchain linking.
- Real-time scheduling and streaming (Rust). Built the scheduler letting many concurrent requests share one model resident on an accelerator: batching, per-request memory-cache management, cancellation, runtime metrics, results streamed incrementally over HTTP and gRPC.
- Distributed ingestion and verification. Designed a protocol verifying how much memory a remote machine actually allocated, using hash-tree commitments so a check costs a fraction of the original work. Rewrote the production ingestion service from Python to Rust/gRPC: bounded worker pools, TLS, health checks, structured telemetry, Docker and CI/CD.
- Front-end and internal tooling (TypeScript, React). Browser-based editor for composing and running processing graphs from data, model and code nodes — node I/O generated from each tool's schema, user Python executed in a Pyodide (WebAssembly) sandbox, runs triggered by webhooks. Also internal CLI and terminal tooling.
GPU / Compiler Engineer, Pyccel · Mohammed VI Polytechnic University (al-Khwarizmi)
Nov 2021 – Apr 2023
Open-source compiler translating scientific Python into C, Fortran and CUDA. Joined as a 6-month intern (Nov 2020 – Apr 2021).
- Built the CUDA backend — GPU kernels and memory management generated from annotated Python — plus core language features including multidimensional arrays and Numba-compatible GPU programming.
- Ported research groups' numerical code to GPU, profiling and rewriting slow paths. Supervised a team of interns on the GPU backend and taught programming labs.
Skills
Languages
Python · C · C++ · Rust · CUDA · TypeScript
ML
PyTorch · Transformers · NumPy · OpenCV · model serving and inference · batching and KV-cache management · embeddings · benchmarking and evaluation
GPU
CUDA · WGSL compute shaders · WebGPU · GPU code generation · profiling · mixed precision · CPU/GPU numerical parity
Publication
Generalized ℒ-Product for High Order Tensors and Applications Using GPU Computations
doi.org/10.1007/978-3-031-89041-3_6
Education
1337 Coding School (42 Network) · Digital Architect2018–2021
FST Tangier · two years of university study2016–2018
Baccalaureate in Physics · Larache2016
Arabic (native) · English (professional, daily working language) · French (professional)