Concepts

Concurrency

Rayzor provides OS threads, channels, and shared state. The compiler checks thread-safety rules: values moved across thread boundaries must be Send; every value shared by reference must also be Sync. Declare these markers with the annotation below; the compiler validates them.

@:derive([Send, Sync]) class SharedData { public var counter:Int; public function new() { counter = 0; } }

Threads

Spawn a closure on an OS thread, then join the returned handle to get its result. The closure is moved into the new thread, and every captured variable must be Send.

var msg = new Message("hello"); var handle = Thread.spawn(() -> { trace(msg.data); return 42; }); var result = handle.join(); // 42 handle.isFinished(); // non-blocking check

Parker provides park/unpark operations: register a thread, block it, and wake it by id. If unpark is called before park, the next park returns immediately.

Channels

Channels support multiple senders and receivers. Capacity zero creates an unbounded channel; a positive capacity creates a bounded channel whose send operation blocks when full. trySend and tryReceive return without blocking.

var ch = new Channel<Message>(10); // bounded Thread.spawn(() -> ch.send(new Message(42))); var msg = ch.receive(); // blocks var maybe = ch.tryReceive(); // null if empty

Select

Use Select to receive from several channels. Select.recv blocks until a channel has a value or is closed; Select.tryRecv polls once and reports index == -1 when nothing was ready.

var r = Select.recv([ch1, ch2]); if (r.index == 1) trace(r.value.v); // 42

A closed, empty channel returns its index and a null value. Check both to detect closure. All channels in one call must share an element type; use Channel<Dynamic> to select across different value types.

Shared state

Arc provides shared ownership across threads. Cloning increments the reference count. Wrap the value in a Mutex when you need to mutate it.

var counter = new Arc(new Mutex(new Counter())); var handle = counter.clone(); // refcount bump Thread.spawn(() -> { var guard = handle.get().lock(); guard.get().value += 1; guard.unlock(); });

The inner type must be Send + Sync to cross a thread boundary. Otherwise, the compiler rejects the spawn. Atomic provides atomic operations for cases that do not need a mutex.

Data-parallel work

WorkerPool distributes work over an index range, such as matrix rows, convolution tiles, or attention heads.

var pool = WorkerPool.global(); pool.parallelFor(1000000, (idx, node) -> { // one worker per NUMA node when nodeCount > 1 });
Multi-node One worker spawned per NUMA node, each pinned before the closure runs
Single-node No fanout. Runs inline on the calling thread. withForcedNodes(N) fans out anyway
Small work Fewer items than twice the node count also runs inline; the fanout would cost more than it saves

SpinPool: reuse workers across calls

For frequently called kernels, creating and joining threads on every call can cost more than the work itself. A SpinPool spawns its workers once and re-dispatches through a lock-free protocol, so a dispatch costs a few atomic stores.

Workers claim chunks of rows from a shared atomic cursor until all rows are processed. With fixed row assignments, an E-core can take three to four times longer than a P-core, holding up the whole call. Chunk stealing lets faster cores take more work without needing to know the CPU topology.

Results stay bit-identical. Each row is computed by exactly one worker with an unchanged per-row reduction order, so a parallel run matches the serial loop exactly.

Topology & affinity

CpuTopology provides the topology information used by WorkerPool. Use it directly when you need explicit thread affinity. The topology is queried once per process, on first use.

Multi-node servers

On multi-socket Linux and Windows, one worker is pinned per NUMA node so allocations land first-touch on that node's memory controller.

UMA hardware and wasm

Node count is 1, every CPU maps to node 0, and binding succeeds as a soft affinity hint. Work runs inline unless you force fanout.

The concurrent package

Thread<T> spawn a closure, join for its result, isFinished for a non-blocking check
Channel<T> MPMC queue: send, receive, trySend, tryReceive, close. Unbounded at capacity 0
Select recv and tryRecv across an array of channels, returning (index, value)
Mutex<T> Exclusive access to the inner value; lock returns a guard, unlock releases it
Arc<T> Atomic refcounted shared ownership; clone bumps the count, not the value
Future<T> Lazy. Nothing runs until await() blocks or then(cb) resolves on a worker
WorkerPool parallelFor and parallelRows over an index range, with NUMA pinning on multi-node systems
SpinPool Persistent workers with chunk-stealing dispatch, for kernels called hundreds of times
Parker registerParkable, park, unpark. The wake primitive, with no lost-wakeup window
CpuTopology multiNode, nodeCount, cpu count and bindToNode for explicit affinity
Runtime implementation

Drop behavior for runtime-managed types, the closure ABI, and the tier ladder.

Architecture →