Architecture
Rayzor adds a tiered runtime and a shared optimization pipeline to Haxe. The native backends generate machine code directly; the WebAssembly backend emits WASM. In tiered execution, code starts immediately in the interpreter; functions are compiled as they get hot.
The supported output targets are native code and WebAssembly. There is no planned JavaScript target.
The pipeline
HIR preserves source-level structure for diagnostics and ownership analysis. MIR represents typed operations in basic blocks, which lets the backends share optimization passes. Type checking happens during AST lowering.
MIR
MIR uses SSA basic blocks with phi nodes and retains type metadata. Collections use BTreeMap to keep iteration order deterministic and code generation reproducible.
MIR prints in an LLVM-like textual form — registers $N, blocks bbN — sorted by id so dumps diff cleanly. RAYZOR_DUMP_MIR=1 prints MIR before optimization; use rayzor dump --diff to inspect changes made by the passes.
Optimization levels
The default is O2. InsertFree runs at every level because it inserts the cleanup required for correct memory management.
O0 still runs required passes. Haxe requires inlining for functions marked inline. Inlining small constructors also exposes the Alloc+GEP pattern that SRA needs to remove allocations. Without it, allocations inside loops cannot be scalarised and leak.
Pass order matters. Inlining exposes Alloc+GEP for SRA. GlobalLoadCaching deduplicates metadata loads so BCE can remove bounds checks. BCE exposes invariant loads for LICM to hoist, preparing loops for unrolling and vectorization.
Tiered execution
Tier 0 is the interpreter, tiers 1–3 use Cranelift, and tier 4 uses LLVM. A function can skip tiers if its call count passes several thresholds at once.
The promotion barrier waits for active JIT calls to finish before replacing a function pointer. The promoter requests a safepoint, waits for the execution counter to reach zero, and swaps the pointer under a write lock. A completed compilation is discarded if the function has already reached a higher tier.
SIMD
SIMD vectors are built into the type system. The rayzor.SIMD* types are @:coreType abstracts that lower directly to MIR vector types. They support float and integer lanes in 128- and 256-bit widths, subject to backend support.
The 256-bit types are not portable. SIMD8i32 and SIMD32i8 are LLVM-only — wasm's v128 and Cranelift have no 256-bit vector type and refuse them rather than narrowing. Gate behind #if llvm with a 128-bit fallback.
The rayzor package
The rayzor package adds memory primitives, native collections specialized by element type, and concurrency APIs checked by the compiler through Send and Sync.
Pointer-sized abstracts that carry machine addresses and are never truncated.
Vec<T> is monomorphized per element type — VecI32, VecF64, packed VecBool, VecPtr — so primitives are contiguous and unboxed.
Message passing and shared state, validated at compile time through Send and Sync.
Windowing, compile-time feature detection, and a built-in spec runner.
WorkerPool uses the detected NUMA topology. On multi-NUMA-node systems the pool pins one worker per node so allocations land first-touch on that node's controller; on UMA hardware and wasm it runs inline on the calling thread unless you force fanout.
Backends
AOT defaults to LLVM; without the llvm-backend feature rayzor aot reports an error.
Memory & object layout
Cleanup is decided at compile time. The HIR-level drop-point analyzer computes last use per variable, tracking loop position, reassignment, block depth, and two escape sets — general escapes and lambda captures, since a captured variable is owned by the closure.
There is no vtable pointer in the header — dispatch resolves the vtable from the class id in slot 0, and interface values are fat pointers wrapped at the new site. Allocation size takes the maximum over the whole extends chain at the allocation site, because an imported parent's fields may not have been visible when the subclass was registered.
On-disk formats
One MIR module plus metadata and cached maps.
Pre-resolved stdlib symbols, for fast startup.
All modules, module table, entry point, build info.
Cache invalidation uses three independent keys: a source content hash, the compiler semver, and a content-derived compiler cache ABI id — the last exists because parser or MIR-shape changes do not bump the semver. Cached maps are keyed by name, never by symbol or type id, because ids are reassigned per compilation.
Pass internals, the runtime ABI, closure conventions and the format specs.