Concepts

Architecture

Rayzor adds a tiered runtime and a shared optimization pipeline to Haxe. The native backends generate machine code directly; the WebAssembly backend emits WASM. In tiered execution, code starts immediately in the interpreter; functions are compiled as they get hot.

The supported output targets are native code and WebAssembly. There is no planned JavaScript target.

The pipeline

01 Parser
Preprocessor and conditional compilation, then a recursive-descent parser — with a legacy parser for error recovery
02 AST
Source syntax verbatim, then macro expansion (interpreter, reification, @:build)
03 TAST
Types, symbol table and type table resolved. Type checking is folded into lowering, not a separate phase
04 HIR
Desugared but still structured — ForIn, TryCatch, Switch, lambdas. Ownership analysis and diagnostics read here
05 MIR
SSA over basic blocks with real phi nodes. Monomorphization runs here, after lowering and the stdlib merge
06 Backends
One optimized MIR consumed by the interpreter, Cranelift, LLVM and WASM alike

HIR preserves source-level structure for diagnostics and ownership analysis. MIR represents typed operations in basic blocks, which lets the backends share optimization passes. Type checking happens during AST lowering.

$ rayzor compile main.hx --stage ast|tast|hir|mir|native

MIR

MIR uses SSA basic blocks with phi nodes and retains type metadata. Collections use BTreeMap to keep iteration order deterministic and code generation reproducible.

IrModule functions · globals · types · string_pool · extern_functions
IrFunction signature · cfg · locals · register_types
IrBasicBlock instructions · terminator · phi_nodes · predecessors

MIR prints in an LLVM-like textual form — registers $N, blocks bbN — sorted by id so dumps diff cleanly. RAYZOR_DUMP_MIR=1 prints MIR before optimization; use rayzor dump --diff to inspect changes made by the passes.

Optimization levels

The default is O2. InsertFree runs at every level because it inserts the cleanup required for correct memory management.

O0 Inlining(15), DCE, UnreachableBlockElim, SRA, CopyProp, DCE
O1 Inlining, DCE, Devirtualization, ConstantFolding, CopyProp, UnreachableBlockElim — the only level without SRA
O2 O1 plus SRA, GlobalLoadCaching, BCE, GVN, CSE, LICM, LoopUnrolling, ControlFlowSimplify, DCE (default)
O3 O2 plus LoopVectorization and TailCallOpt

O0 still runs required passes. Haxe requires inlining for functions marked inline. Inlining small constructors also exposes the Alloc+GEP pattern that SRA needs to remove allocations. Without it, allocations inside loops cannot be scalarised and leak.

Pass order matters. Inlining exposes Alloc+GEP for SRA. GlobalLoadCaching deduplicates metadata loads so BCE can remove bounds checks. BCE exposes invariant loads for LICM to hoist, preparing loops for unrolling and vectorization.

Tiered execution

Tier 0 is the interpreter, tiers 1–3 use Cranelift, and tier 4 uses LLVM. A function can skip tiers if its call count passes several thresholds at once.

T0 Interpreted MIR interpreter at O0 — register-based, since MIR is already SSA. Instant startup
T1 Baseline Cranelift, no opt, O0. Compiled inline on the main thread
T2 Standard Cranelift speed, O1. Routed to a background broker thread
T3 Optimized Cranelift speed, O2. Background, one adapter and bead registry per tier
T4 Maximum LLVM at O3, queued and drained on the main thread — add_global_mapping requires it

The promotion barrier waits for active JIT calls to finish before replacing a function pointer. The promoter requests a safepoint, waits for the execution counter to reach zero, and swaps the pointer under a write lock. A completed compilation is discarded if the function has already reached a higher tier.

SIMD

SIMD vectors are built into the type system. The rayzor.SIMD* types are @:coreType abstracts that lower directly to MIR vector types. They support float and integer lanes in 128- and 256-bit widths, subject to backend support.

SIMD4f
4 × f32
128-bit
SIMD4i32
4 × i32
128-bit
SIMD16i8
16 × i8
128-bit
SIMD8i32
8 × i32
256-bit
SIMD32i8
32 × i8
256-bit
import rayzor.SIMD4f; var a:SIMD4f = (1.0, 2.0, 3.0, 4.0); // tuple literal var b = SIMD4f.splat(2.0); // broadcast to 4 lanes var c = a * b + a; // @:op overloads var d = a.dot(b); // dot, sum, normalize, magnitude, lerp
Representation Decided from the type's identity ahead of the underlying type, so a vector is never truncated to a 64-bit param at a call boundary
Construction Tuple literal, array literal via @:from, SIMD4f.make(x, y, z, w), or SIMD4f.splat(v) to broadcast one scalar across all four lanes
Operations Arithmetic operators, lane access with array syntax, and float methods: dot, sum, sqrt, abs, min, max, rounding, normalize, magnitude, lerp
Quantized kernels The integer types carry widening dot-accumulate — dot, dotI8I7, dotI8U8 — which lower to a single VNNI instruction on x86-64
Native paths Vectorized f32 CPU paths for NEON and SSE2, with a scalar fallback where neither is available
WASM SIMD128 The WebAssembly backend carries the same vector work through SIMD128
Tier policy The interpreter executes vector types, and functions that use SIMD are promoted to Baseline on first call so vector work runs compiled from the start
Vectorization LoopVectorization runs at O3, on the clean loop bodies LICM leaves behind

The 256-bit types are not portable. SIMD8i32 and SIMD32i8 are LLVM-only — wasm's v128 and Cranelift have no 256-bit vector type and refuse them rather than narrowing. Gate behind #if llvm with a 128-bit fallback.

The rayzor package

The rayzor package adds memory primitives, native collections specialized by element type, and concurrency APIs checked by the compiler through Send and Sync.

Memory & systems

Pointer-sized abstracts that carry machine addresses and are never truncated.

Ptr Ref Box Usize Mem Bytes Slice CString Double Atomic
Collections

Vec<T> is monomorphized per element type — VecI32, VecF64, packed VecBool, VecPtr — so primitives are contiguous and unboxed.

Vec<T> Result<T,E>
Concurrency

Message passing and shared state, validated at compile time through Send and Sync.

Thread Channel<T> Select Mutex Arc Future WorkerPool SpinPool Parker CpuTopology
Concurrency guide →
Platform

Windowing, compile-time feature detection, and a built-in spec runner.

Window Key EventType WindowStyle CC Spec
var pool = WorkerPool.global(); pool.parallelFor(1000000, (idx, node) -> { ... }); // one worker per NUMA node var f = Future.create(() -> heavy()); // lazy — nothing runs yet f.then(v -> trace(v)); // or f.await() to block

WorkerPool uses the detected NUMA topology. On multi-NUMA-node systems the pool pins one worker per node so allocations land first-touch on that node's controller; on UMA hardware and wasm it runs inline on the calling thread unless you force fanout.

Backends

interpreter Register-based MIR interpreter. Instant startup, tier 0
cranelift JIT tiers 1–3. Fast compile; fuses fmul into fma within a block
llvm Default AOT path, top JIT tier, and a whole-module upgrade
wasm MIR → core WASM → WASI P2 component. Linear memory, SIMD128
wgsl @:shader classes → WGSL at compile time. Not a general target

AOT defaults to LLVM; without the llvm-backend feature rayzor aot reports an error.

Memory & object layout

Cleanup is decided at compile time. The HIR-level drop-point analyzer computes last use per variable, tracking loop position, reassignment, block depth, and two escape sets — general escapes and lambda captures, since a captured variable is owned by the closure.

slot 0 __type_id : i64 ← stable name-hash class id slot 1 first user field ... every slot is 8 bytes

There is no vtable pointer in the header — dispatch resolves the vtable from the class id in slot 0, and interface values are fat pointers wrapped at the new site. Allocation size takes the maximum over the whole extends chain at the allocation site, because an imported parent's fields may not have been visible when the subclass was registered.

AutoDrop Compiler emits Free — user classes allocated with new
AutoDropWithDtor Run the user's drop(), then Free — @:derive(Drop)
ManualDrop @:manualDrop; never auto-freed
RuntimeManaged Runtime owns the lifetime — Thread, Channel, Arc, Mutex
NoDrop Primitives, arrays, Dynamic

On-disk formats

.blade BLAD

One MIR module plus metadata and cached maps.

.bsym BSYM

Pre-resolved stdlib symbols, for fast startup.

.rzb RZBF

All modules, module table, entry point, build info.

Cache invalidation uses three independent keys: a source content hash, the compiler semver, and a content-derived compiler cache ABI id — the last exists because parser or MIR-shape changes do not bump the semver. Cached maps are keyed by name, never by symbol or type id, because ids are reassigned per compilation.

The full architecture doc

Pass internals, the runtime ABI, closure conventions and the format specs.

Read it on GitHub ↗