How are answers generated? · Local model loads on first question · WebGPU/WASM inference stays on this device.