AVAILABLE FOR ONE CONTRACT ENGAGEMENT

ANDREAS FALKENBERG — PHD

Compilers forcustom AIsilicon.

I put PyTorch and ONNX models onto bespoke accelerators — LLVM/MLIR lowering, graph partitioning and hardware-software co-design for AI-silicon teams.

RÉSUMÉ — PDF
LLVM//MLIR//PYTORCH//ONNX//CUSTOM BACKENDS//GRAPH PARTITIONING//QUANTIZATION//SYSTEMC//FPGA / HLS//KERNEL CODEGEN//HW-SW CO-DESIGN//

01 // WHAT I DO

Four ways I plug into your stack.

01

LLVM / MLIR Compiler Architecture

End-to-end compiler stacks for custom silicon: dialect design, pass pipelines, scheduling and codegen shaped around your hardware.

DIALECTSPASS PIPELINESCODEGENCOST MODELS
02

PyTorch / ONNX Lowering & Custom Backends

Bridging model graphs to your target: torch-mlir / ONNX lowering paths, custom backend integration, operator coverage and numerics that hold.

TORCH-MLIRONNXBACKEND HOOKSNUMERICS
03

Hardware-Aware Partitioning & Quantization

Making models fit the machine: partitioning across compute tiles, quantization and QAT strategy, memory-aware scheduling and tiling.

PARTITIONINGQUANT / QATTILINGMEMORY PLANNING
04

ASIC / SoC & FPGA, SystemC Co-Design

Working the hardware-software boundary: ISA and intrinsic feedback, SystemC performance models, FPGA bring-up and early compiler validation for SoC teams.

SYSTEMCFPGA BRING-UPISA FEEDBACKPERF MODELING

02 // EXPERIENCE

Named employers. Real hardware.

01

AMD / Xilinx

Ryzen AI

Upstream torch-mlir contributions and custom-backend work lowering ONNX/PyTorch models onto Ryzen AI NPUs — including NPU kernel integration and compiler validation.
02

BrainChip

Sole Compiler Engineer

Built the complete MLIR-based compiler stack for BrainChip's next-generation neuromorphic processor — from ONNX, PyTorch, and TensorFlow graph ingestion through custom dialect design, middle-end optimization passes (fusion, tiling, quantization, layout), custom ISA backend codegen, and runtime integration. Defined the correctness validation strategy and led a researcher to scale it to 800+ test cases.
03

d-Matrix

Inference silicon

Graph-to-hardware mapping for transformer inference on in-memory compute.
04

Metawave

Radar perception

Radar accuracy 69% → 95% — signal-chain ML taken from lab to field.
05

Ostendo

Photonic display

Designed the backend ASIC for a photonic embedded display from first principles through physical tapeout — architecture, RTL, and validation.

* SELECTED ENGAGEMENTS — FULL DETAIL IN THE RÉSUMÉ PDF

2×

PHD — DOCTORATES

52

PATENTS — 33 GRANTED

23

PEER-REVIEWED PUBLICATIONS

5

INVITED TALKS

6

UNIVERSITY COURSES

04 // CONTACT

One engagement.
Let's make it count.

andreas@falkenbergtech.com

PREFER EMAIL? WRITE DIRECTLY — IT LANDS IN MY INBOX.

AVAILABLE FOR ONE CONTRACT ENGAGEMENT