AI Inference #

Neural inference with ONNX #

ONNX Runtime support is optional at build time and opt-in at app level. Add neural to app.json.

app.json
{
  "neural": true
}

Then check capability before loading a model.

main.js
// Check the capability before attempting to load ONNX models.
if (!sys.capabilities.neural.available) {
  // Choose a fallback path for builds without neural inference.
  console.log('No neural runtime in this build');
}

Models must be ONNX data. loadModel(path) reads .onnx files inside the project directory; paths are relative and path traversal is rejected. loadModelFromBuffer(buffer) accepts the same ONNX bytes from an ArrayBuffer / typed-array view.

main.js
// Load an ONNX model from the project directory.
const model = sys.neural.loadModel('models/classifier.onnx');
const fetchedModel = sys.neural.loadModelFromBuffer(modelBytes);
// Inspect model inputs and outputs.
const info = sys.neural.getModelInfo(model);

// Use the first input name declared by the model.
const inputName = info.inputs[0].name;
// Use the first output name declared by the model.
const outputName = info.outputs[0].name;

// Allocate input data for a 1x3x224x224 tensor.
const data = new Float32Array(1 * 3 * 224 * 224);
// Run inference with a named input tensor.
const result = sys.neural.run(model, {
  // Bind the typed array to the model's input name.
  [inputName]: data
});

// Read the model output by name.
const output = result[outputName];

JavaScript uses typed arrays. Lua uses tables with data, shape, and dtype. WebAssembly uses staged host imports that copy tensors into and out of guest memory.

The staged JavaScript form is useful when you want the same conceptual flow as Lua and WebAssembly.

main.js
// Stage input data on the model handle.
sys.neural.setInput(model, inputName, data);
// Run inference using the staged inputs.
sys.neural.run(model);
// Read the staged output tensor by name.
const output = sys.neural.getOutput(model, outputName);

Unload models when a long-lived app no longer needs them. This matters for tools that let users switch models, for games that load scene-specific inference assets, and for Android devices where model memory can be much tighter than on a desktop workstation.

main.js
// Release model resources when they are no longer needed.
sys.neural.unloadModel(model);

LlamaCpp inference #

Platform notes #

ONNX and LlamaCpp inference depend on platform permissions or linked backends. Use capability checks and clear UI states.