Backends_deprecated.Multidev_cc_bval sexp_of_code : code -> Sexplib0.Sexp.tval sexp_of_code_batch : code_batch -> Sexplib0.Sexp.tval empty_optimize_ctx : Base.unit -> Ir.Low_level.optimize_ctxval get_optimize_ctx : code -> Ir.Low_level.optimize_ctxval get_optimize_ctx_batch : code_batch -> Ir.Low_level.optimize_ctxval compile :
Ir.Low_level.optimize_ctx ->
?name:Base.string ->
?lowered_transform:(Ir.Low_level.optimized -> Ir.Low_level.optimized) ->
?lowered_transforms:
(Ir.Low_level.optimized -> Ir.Low_level.optimized Base.list) ->
Ir.Indexing.unit_bindings ->
Ir.Assignments.comp ->
codename is used to derive names for compilation artifacts. If omitted, it's derived via Assignments.get_name_exn. lowered_transform is applied to the optimized lowered code before backend compilation — the seam where schedule transforms (and hand-annotating tests) rewrite loops with hardware axis types, barriers and shared placements (docs/proposals/axis-types-for-loops.md). lowered_transforms is the plural variant for transforms that split the routine into several kernels (fission, Schedule.fission_scheduled): the returned segments compile as one fissioned routine and run back-to-back on the routine's stream with a device-side event chained at each boundary, exactly as Schedule.maybe_default_schedules' segments do. It must return a non-empty list; passing both transforms raises Invalid_argument.
val compile_batch :
Ir.Low_level.optimize_ctx ->
?names:Base.string Base.array ->
?occupancy:(name:Base.string -> src_n:Base.int -> Base.bool) ->
Ir.Indexing.unit_bindings ->
Ir.Assignments.comp Base.array ->
code_batchcompile_batch vs. compile is mostly about improving the compile time and debugging convenience by generating fewer files -- ideally does not affect execution, but there can be backend-specific differences. Only array entries for which occupancy returns true are included. names are used to derive names for compilation artifacts. If omitted, they're derived via Assignments.get_name_exn.
include Ir.Backend_intf.Backend_device_commoninclude Ir.Backend_intf.Deviceinclude Ir.Backend_intf.Device_typesinclude Ir.Backend_intf.Device_config_commonval sexp_of_dev : dev -> Sexplib0.Sexp.tval sexp_of_runner : runner -> Sexplib0.Sexp.tAn event tracks if a device's runner finished computing past a particular point in its schedule. These values are used internally for scheduling across devices/queues of the backend, and can be used for explicit scheduling.
val sexp_of_event : event -> Sexplib0.Sexp.ttype nonrec device = (dev, runner, event) Ir.Backend_intf.deviceval sexp_of_device : device -> Sexplib0.Sexp.ttype nonrec context = (dev, runner, event) Ir.Backend_intf.contextval sexp_of_context : context -> Sexplib0.Sexp.tinclude Ir.Backend_intf.Slab_alloc with type device := deviceval alloc_pool :
?mode:Ir.Tnode.memory_mode ->
device ->
pool_id:Base.int ->
size_in_bytes:Base.int ->
alignment:Base.int ->
Base.unitAllocates the slab for pool_id on device. The optional ?mode carries the tnode's memory mode so backends can pick a storage mode (Metal private vs. shared); backends that do not care ignore it.
val free_pool : (device -> pool_id:Base.int -> Base.unit) Base.optionFrees the slab for pool_id and drops its table entry. None for backends that rely on GC.
val memset_zero :
device ->
pool_id:Base.int ->
offset:Base.int ->
size_in_bytes:Base.int ->
Base.unitZero-initializes size_in_bytes at base_of pool_id + offset.
val make_context :
?ctx_buffers:Ir.Backend_intf.ctx_buffers ->
?optimize_ctx:Ir.Low_level.optimize_ctx ->
device ->
contextReturns a context without a parent.
val make_child :
?ctx_buffers:Ir.Backend_intf.ctx_buffers ->
?optimize_ctx:Ir.Low_level.optimize_ctx ->
?merge_buffer_node:Ir.Tnode.t Base.option ->
context ->
contextReturns a context with the same Backend_intf.context.device, Backend_intf.context.ctx_buffers, Backend_intf.context.optimize_ctx, Backend_intf.context.merge_buffer_node if omitted, as the given context's, which is also the Backend_intf.context.parent.
val get_name : device -> Base.stringval sync : event -> Base.unitBlocks till the event completes, if it's not done already.
It is rarely needed to call sync explicitly, because it should always be called internally when necessary, in particular before extracting values from host.
val is_done : event -> Base.boolWhether the event completed.
Schedules waiting for the given event on the context's device.
NOTE: it should rarely be needed to call will_wait_for explicitly, because it should always be called internally when necessary.
Returns a sexp description of the properties of all devices. A function so that computing it (device enumeration) does not run at backend-module initialization: singleton backends instantiate eagerly at program startup, where touching a driver could fail runs that never use the backend.
val hardware_limits : Base.unit -> Ir.Backend_intf.hardware_limitsConservative per-workgroup device limits: on multi-device backends the minimum across the devices, so code compiled once (compilation is not per-device) is valid wherever it links. All-None for backends that do not bind hardware axes. A function for the same reason as static_properties: computing it (device enumeration) must not run at backend-module initialization.
val get_used_memory : device -> Base.intReturns (an upper bound of) the memory used for arrays, in bytes.
Global debug information; backend-specific and might evolve independently on the backends.
val get_debug_info : device -> Base.Sexp.tPer-device debug information; backend-specific and might evolve independently on the backends
val await : device -> Base.unitBlocks till the device becomes idle, i.e. synchronizes the device's runner.
Returns the event indicating if any currently running or scheduled computations on the device have completed.
val is_idle : device -> Base.boolWhether the device's runner is currently waiting for work.
val get_device : ordinal:Base.int -> deviceval link : context -> code -> context Ir.Backend_intf.routineReturns the routine for the code's procedure, in a new context derived from the given context.
val link_batch :
context ->
code_batch ->
context * context Ir.Backend_intf.routine Base.option Base.arrayReturns the routines for the procedures included in the code batch. The returned context is downstream of all the returned routines.
include Ir.Backend_intf.With_buffer_retrieval_and_syncing
with type device := device
and type context := context
and type event := eventval from_host : context -> Ir.Tnode.t -> Ir.Ndarray.t -> Base.boolfrom_host ctx tn src schedules a copy of the explicit host buffer src into tn's in-context device buffer and returns true, or returns false if the node is not in context. After gh-ocannl-333 the host buffer is supplied by the caller (e.g. Context.set_values); it is no longer read from the tensor node.
val init_from_host : context -> Ir.Tnode.t -> Ir.Ndarray.t -> contextSchedules a copy from the explicit host buffer to context: a variant of from_host that requires the input context to not contain the tensor node, and outputs the context with the tensor node.
val to_host : context -> Ir.Tnode.t -> Ir.Ndarray.t -> Base.boolto_host ctx tn dst schedules a copy of tn's in-context device buffer into the explicit host buffer dst and returns true, or returns false if the node is not in context. After gh-ocannl-333 the destination buffer is supplied by the caller (e.g. Context.to_host); it is no longer the tensor node's own array.
val device_to_device :
Ir.Tnode.t ->
into_merge_buffer:Ir.Backend_intf.merge_buffer_use ->
dst:context ->
src:context ->
context Ir.Backend_intf.routine Base.optiondevice_to_device tn ~into_merge_buffer ~dst ~src builds a transfer routine instead of scheduling the copy directly. The caller schedules it (e.g. via Task.run r.schedule) or links a consumer against r.context. It returns:
None if there is nothing to transfer: the node is absent from src; or, for into_merge_buffer=No, the node is absent from dst or the source and destination buffers are physically the same.Some r otherwise. Running r.schedule waits for writing into the tensor node on src to finish, then performs the copy and updates the writer event.into_merge_buffer=No, the copy goes from src to dst; r.context is a child of dst inheriting its Backend_intf.context.merge_buffer_node.into_merge_buffer=Copy, the copy goes from src to the merge buffer of dst's stream; r.context is a child of dst with merge_buffer_node = Some tn, so that linking a consumer of the merge buffer against r.context statically verifies the node.val init_from_device : Ir.Tnode.t -> dst:context -> src:context -> contextSchedules a copy from src to dst: a variant of device_to_device with into_merge_buffer=No that requires the input src context to not contain the tensor node, and outputs the dst context with the tensor node.
val sync_device : device -> Base.unitSynchronizes all the streams on a device, and cleans up (removes) all associated events.