Concurrency
Rayzor provides OS threads, channels, and shared state. The compiler checks thread-safety rules: values moved across thread boundaries must be Send; every value shared by reference must also be Sync. Declare these markers with the annotation below; the compiler validates them.
Threads
Spawn a closure on an OS thread, then join the returned handle to get its result. The closure is moved into the new thread, and every captured variable must be Send.
Parker provides park/unpark operations: register a thread, block it, and wake it by id. If unpark is called before park, the next park returns immediately.
Channels
Channels support multiple senders and receivers. Capacity zero creates an unbounded channel; a positive capacity creates a bounded channel whose send operation blocks when full. trySend and tryReceive return without blocking.
Select
Use Select to receive from several channels. Select.recv blocks until a channel has a value or is closed; Select.tryRecv polls once and reports index == -1 when nothing was ready.
A closed, empty channel returns its index and a null value. Check both to detect closure. All channels in one call must share an element type; use Channel<Dynamic> to select across different value types.
Shared state
Arc provides shared ownership across threads. Cloning increments the reference count. Wrap the value in a Mutex when you need to mutate it.
The inner type must be Send + Sync to cross a thread boundary. Otherwise, the compiler rejects the spawn. Atomic provides atomic operations for cases that do not need a mutex.
Data-parallel work
WorkerPool distributes work over an index range, such as matrix rows, convolution tiles, or attention heads.
SpinPool: reuse workers across calls
For frequently called kernels, creating and joining threads on every call can cost more than the work itself. A SpinPool spawns its workers once and re-dispatches through a lock-free protocol, so a dispatch costs a few atomic stores.
Workers claim chunks of rows from a shared atomic cursor until all rows are processed. With fixed row assignments, an E-core can take three to four times longer than a P-core, holding up the whole call. Chunk stealing lets faster cores take more work without needing to know the CPU topology.
Topology & affinity
CpuTopology provides the topology information used by WorkerPool. Use it directly when you need explicit thread affinity. The topology is queried once per process, on first use.
On multi-socket Linux and Windows, one worker is pinned per NUMA node so allocations land first-touch on that node's memory controller.
Node count is 1, every CPU maps to node 0, and binding succeeds as a soft affinity hint. Work runs inline unless you force fanout.
The concurrent package
Drop behavior for runtime-managed types, the closure ABI, and the tier ladder.