How I got a multilingual PII redaction model running on Akamai Functions (Fermyon Spin + WebAssembly), after the obvious approaches failed one by one.
The goal
I wanted an HTTP endpoint that takes text and returns it with personal data replaced by placeholders:
curl -s -X POST https://<app>.fwf.app/redact \
-H "Content-Type: application/json" \
-d '{"text": "My name is John Doe, I live in Paris, and my email is john.doe@example.com."}'
{
"redactedText": "My name is [GIVEN_NAME_1] [SURNAME_1], I live in [CITY_1], and my email is [EMAIL_1].",
"items": [
{ "label": "GIVEN_NAME", "original": "John", "placeholder": "[GIVEN_NAME_1]", "start": 11, "end": 15 },
{ "label": "SURNAME", "original": "Doe", "placeholder": "[SURNAME_1]", "start": 16, "end": 19 },
{ "label": "CITY", "original": "Paris", "placeholder": "[CITY_1]", "start": 31, "end": 36 },
{ "label": "EMAIL", "original": "john.doe@example.com", "placeholder": "[EMAIL_1]", "start": 54, "end": 74 }
]
}
The model is Desert Ant Labs’ redact: a multilingual PII token classifier.
| Property | Value |
|---|---|
| Architecture | BERT-style encoder (XLM-RoBERTa / MiniLM lineage) |
| Parameters | ~23M, int8 weights |
| Layers / hidden / heads / FFN | 6 / 384 / 12 / 1536 |
| Sequence length | fixed 256 tokens |
| Output | 89 BIOES labels over 20 PII categories (GIVEN_NAME, SURNAME, CITY, EMAIL, CREDIT_CARD, IBAN, …) |
| Distribution | npm package @desert-ant-labs/redact and redact.tflite on Hugging Face |
And I wanted it on Akamai Functions: WebAssembly components running on Spin, close to users, with no containers, no GPUs and no sidecars.
This post covers how I got there, including the dead ends. The follow-up post covers how I then took latency from ~1 s down to ~260 ms.
Understanding the platform first
Three properties of Akamai Functions / Spin decide almost everything below:
- Code runs as a WebAssembly component in a WASI sandbox. No native libraries, no
dlopen, no C FFI to the host. - JavaScript runs inside StarlingMonkey (SpiderMonkey compiled to Wasm, via ComponentizeJS). There is no DOM, no Web Workers, and — the one that bit us — no nested
WebAssemblyAPI. - A fresh instance is created for every request. Whatever you set up at start-up, you pay for on every request.
I didn’t fully understand all three at the start. I learned them in order.
Attempt 1: “Just use the npm package” (JavaScript)
The model ships as an npm package, and Spin has a JavaScript template. The obvious first try:
import { Redact } from "@desert-ant-labs/redact";
let model = null;
export async function handleRequest(request) {
if (!model) model = await Redact.load();
const { text } = await request.json();
const result = await model.redaction(text);
return new Response(JSON.stringify(result), {
headers: { "content-type": "application/json" },
});
}
Problem 1: dynamic Node imports
Componentising (j2w) failed straight away. The package’s default (Node) entry point did:
const fs = await import("node:fs");
ComponentizeJS rejects that:
Dynamic module import is disabled or not supported in this context
The Node build also loads the model through koffi, a native FFI library that calls into .dylib/.so files. That can never work inside a Wasm sandbox.
Problem 2: the browser build hangs silently
So I pointed the bundler at the browser build (conditions: ['browser']). It compiled. It deployed. Every request failed with:
ERROR: no fetch-event handler triggered
No stack trace, no exception. Reading @desert-ant-labs/redact/browser.js explained it. At the top level of the module:
const sdk = await createWasmSdk(/* ... */); // top-level await
createWasmSdk uses LiteRT.js, which loads its own Wasm runtime by injecting <script> tags and compiling Wasm with the WebAssembly API. Inside StarlingMonkey:
typeof WebAssembly === "undefined" // true
So the top-level promise never resolved, the module never finished initialising, and addEventListener('fetch', ...) was never registered. Spin saw a component with no handler.
Lesson: JavaScript ML packages that depend on nested WebAssembly, DOM script injection, Web Workers or native FFI cannot run in Spin’s JS runtime. The JS project (
redact-akamai-function/) was left as a stub, and I moved to Rust.
Attempt 2: an off-the-shelf Rust runtime (tract-tflite)
Rust compiles straight to wasm32-wasip1, and tract is a pure-Rust inference engine with a TFLite frontend. I downloaded redact.tflite and tried:
let model = tract_tflite::tflite()
.model_for_path("redact.tflite")?
.into_optimized()?
.into_runnable()?;
It failed on load:
Error: Translating proto model to model
Caused by: Error in TF file for operator Operator { ... }. No prior computation nor constant for input 115
Tensor 115 was the word embedding table. Why would a constant be “missing”?
Digging into the flatbuffer
I inspected the file with the Python tflite package:
import tflite
buf = open("redact.tflite", "rb").read()
model = tflite.Model.GetRootAsModel(buf, 0)
sub = model.Subgraphs(0)
for i in range(sub.TensorsLength()):
t = sub.Tensors(i)
b = model.Buffers(t.Buffer())
print(i, t.Name().decode(), t.Type(), t.ShapeAsNumpy(),
"inline:", b.DataLength(), "offset:", b.Offset(), "size:", b.Size())
101 of the 122 buffers had DataLength() == 0. Their data wasn’t inside the flatbuffer at all. They used TFLite’s external buffer offset extension: Buffer.offset and Buffer.size point to raw bytes appended after the end of the flatbuffer. That’s how large models get past flatbuffers’ 2 GB limit, and some exporters use it for all big tensors. tract-tflite only reads inline buffers, so to it every weight was missing.
Even if tract had loaded the model, there was a second, deeper problem: Spin creates a new instance per request. Any runtime that parses a model graph, allocates tensors and optimises the plan at start-up does that on every request. For graph-based runtimes (tract, ONNX Runtime, the TFLite interpreter compiled to Wasm) that is typically hundreds of milliseconds to seconds before any real work starts.
Lesson: On a per-request-instance platform, start-up cost is per-request cost. I needed a format that needs no parsing at all.
The design that worked: flat weights + hand-written forward pass
The plan:
- Extract every weight from the
.tfliteinto a single flat binary file in a fixed, known order. - Embed that file into the Wasm binary with
include_bytes!, so the weights are in memory as soon as the instance starts. - Write the forward pass by hand in Rust — it’s only six BERT layers.
- Verify against the official LiteRT runtime until the outputs match.
redact.tflite ──extract.py──▶ assets/model.bin ──include_bytes!──▶ Wasm component
│
Model::load (slices, <1 ms)
│
forward(): 6 layers in Rust + Wasm SIMD
Step 1: reverse-engineering the graph (extract.py)
There’s no metadata that says “this is the Q projection of layer 3”. I had to recover the architecture from the graph’s operators.
Reading external buffers
First, a raw() helper that understands both inline and external buffers:
def raw(idx):
"""Bytes of a constant tensor, honouring the external-buffer extension."""
t = sub.Tensors(idx)
b = model.Buffers(t.Buffer())
off = b.Offset()
if off and off > 1:
return buf[off:off + b.Size()] # data lives past the flatbuffer
assert b.DataLength() > 0, f"tensor {idx} is not constant"
return b.DataAsNumpy().tobytes()
Finding the matrix multiplies
A BERT layer has six dense layers: Q, K, V, output projection, FFN-in, FFN-out. With 6 layers plus the classifier I expected 37 FULLY_CONNECTED ops, and found exactly 37. In topological order they come out as [Q, K, V, O, FF1, FF2] × 6 + classifier:
fcs = []
for k in range(sub.OperatorsLength()):
op = sub.Operators(k)
if opnames[op.OpcodeIndex()] == "FULLY_CONNECTED":
ins = [int(x) for x in op.InputsAsNumpy()]
fcs.append((ins[1], ins[2])) # (weights, bias)
assert len(fcs) == 37, len(fcs)
Checking the quantisation scheme
Each weight tensor carries its own quantisation info. I asserted the assumptions rather than hoping:
def scales(idx):
q = sub.Tensors(idx).Quantization()
s = np.array([q.Scale(i) for i in range(q.ScaleLength())], dtype=np.float32)
assert q.QuantizedDimension() == 0, "expected per-output-channel quantisation"
zp = np.array([q.ZeroPoint(i) for i in range(q.ZeroPointLength())])
assert (zp == 0).all(), "expected symmetric quantisation"
return s
So: int8 weights, one float scale per output row, zero-point 0.
Finding LayerNorm
TFLite had no LayerNorm op here; it was decomposed into mean / subtract / variance / rsqrt / MUL / ADD. The affine part (γ, β) shows up as a MUL by a constant [384] vector followed by an ADD of another. I found 12 such pairs: 1 after the embeddings, and 2 per layer except the very last one.
Folded embeddings
Position embeddings and token-type embeddings were two constant [1, 256, 384] tensors added after the word-embedding lookup. Since the sequence length is fixed and token type is always 0, I folded them into one bias:
put_f32(f32(bias_consts[0]).reshape(256, 384) + f32(bias_consts[1]).reshape(256, 384))
The final layout
word_emb int8 [31475*384] + word_scale f32 [31475]
emb_bias f32 [256*384] (position + token_type, folded)
emb_ln_gamma f32 [384] + emb_ln_beta f32 [384]
6 x layer:
q,k,v,o : w int8[384*384] + scale f32[384] + bias f32[384]
attn_ln : gamma f32[384] + beta f32[384]
ffn1 : w int8[1536*384] + scale f32[1536] + bias f32[1536]
ffn2 : w int8[384*1536] + scale f32[384] + bias f32[384]
ffn_ln : gamma f32[384] + beta f32[384]
cls int8 [89*384] + scale f32[89] + bias f32[89]
The script computes the expected size and checks it — 23,463,060 bytes. That check caught an off-by-one early on:
print(f"wrote {DST}: {len(out):,} bytes (expected {expected:,}) match={len(out)==expected}")
Step 2: loading without parsing
On the Rust side, “loading” is just walking a cursor over the embedded bytes. Weight matrices are slices into the binary, not copies:
static MODEL_BLOB: &[u8] = include_bytes!("../assets/model.bin");
struct Cursor<'a> { b: &'a [u8], at: usize }
impl<'a> Cursor<'a> {
fn i8s(&mut self, n: usize) -> &'a [i8] {
let s = &self.b[self.at..self.at + n];
self.at += n;
// i8 and u8 have identical size and alignment.
unsafe { core::slice::from_raw_parts(s.as_ptr() as *const i8, n) }
}
fn qmat(&mut self, out: usize, inp: usize) -> QMat<'a> {
let w = self.i8s(out * inp);
let scale = self.f32s(out);
let bias = self.f32s(out);
QMat { w, scale, bias, out, inp }
}
}
impl<'a> Model<'a> {
pub fn load(blob: &'a [u8]) -> Model<'a> {
let mut c = Cursor { b: blob, at: 0 };
let emb = c.i8s(VOCAB * D);
// ... embeddings, 6 layers, classifier ...
assert_eq!(c.at, blob.len(), "model.bin size mismatch");
Model { /* ... */ }
}
}
Model::load takes under a millisecond. The 22 MB of weights cost nothing at start-up because they are already part of the component’s memory image.
Step 3: a golden reference, and three surprises
Before trusting any Rust output, I set up a reference: the official LiteRT Python runtime (ai-edge-litert) running redact.tflite on the benchmark sentence, saving input ids and logits.
np.save("golden_logits.npy", reference_logits)
json.dump({"ids": input_ids, "text": text}, open("golden_input.json", "w"))
Then I compared our Rust forward pass against it. Three surprises came out of that.
Surprise A: the attention mask must be all ones
I started with the standard Hugging Face convention: attention_mask = 1 for real tokens, 0 for padding. The model found the email address — and nothing else. No John, no Doe, no Paris.
Setting the mask to 1 everywhere, including padding, restored every detection and matched LiteRT. The exported graph was traced with an all-ones mask, and that’s the behaviour it was calibrated on.
Surprise B: padding is load-bearing
Since the mask is all ones, the model attends to padding tokens. So the padding itself affects the output:
- Pad token
0instead of1→ logits shifted. - Trimming the sequence below 256 → logits shifted by up to 13.28.
So every request runs a full 256-token window padded with id 1:
pub const PAD_ID: u32 = 1; // The graph is fixed at 256 tokens and attends over the whole window, so // longer input is processed as successive full windows. let mut window = vec![PAD_ID; SEQ]; window[..chunk_end - chunk_start].copy_from_slice(&ids[chunk_start..chunk_end]); let logits = model.forward(&window, chunk_end - chunk_start);
Surprise C: float is “too accurate”
Our first matrix multiply dequantised the int8 weights and multiplied in f32. Mathematically that’s more precise than int8 — and it agreed with LiteRT on only 86% of token labels. The error grew layer by layer.
The reason is TFLite’s hybrid quantisation. For an int8-weight FULLY_CONNECTED with float input, TFLite:
- Quantises each input row to int8 on the fly:
scale = max|x| / 127,q = round(x / scale). - Does an integer dot product with the int8 weights.
- Dequantises:
y = bias + acc × scale_in × scale_w.
The model was calibrated with that rounding in the loop. Reproducing it exactly:
/// TFLite hybrid FullyConnected.
///
/// Emulating the dynamic activation quantisation matters: a plain f32
/// matmul is *more precise* than the reference and drifts from it by about
/// one quantisation step per layer.
fn quantize_row(xr: &[f32], q: &mut [i8]) -> f32 {
let mut amax = 0f32;
for &a in xr {
let m = if a < 0.0 { -a } else { a };
if m > amax { amax = m; }
}
if amax == 0.0 { q.fill(0); return 0.0; }
let inv = 127.0 / amax;
for k in 0..xr.len() {
let v = (xr[k] * inv).round();
q[k] = if v > 127.0 { 127 } else if v < -127.0 { -127 } else { v as i8 };
}
amax / 127.0
}
// ... then per output o:
y[o] = bias[o] + (acc as f32) * scale_in * scale_w[o];
Result: 100% label agreement with LiteRT, max logit difference 0.31.
Bonus surprise: folded LayerNorm affine
Tracing the graph more closely showed that the LayerNorm γ/β were not applied before Q/K/V and the FFN. They had been folded into those layers’ weights. The graph had two versions of each normalised tensor:
- Pre-affine (just normalised) → feeds Q/K/V or the FFN.
- Post-affine (× γ + β) → feeds the residual connection.
So the layer structure is:
layer_norm(&x, n, &mut norm); // pre-affine apply_affine(&norm, n, &ly.attn_gamma, &ly.attn_beta, &mut affine); // post-affine ly.ff1.matmul(&norm, ...); // FFN consumes the pre-affine tensor // ... x[i] = proj[i] + affine[i]; // residual adds the post-affine tensor
The final layer’s LayerNorm has an identity affine (γ=1, β=0), which extract.py writes explicitly.
Step 4: making it fast enough to be usable
The first correct version was scalar Rust, and a request took ~3.2 s. I tried relying on LLVM’s auto-vectoriser with SIMD enabled; the int8 dot product barely improved.
Enabling Wasm SIMD is one line in .cargo/config.toml:
[target.wasm32-wasip1] rustflags = ["-C", "target-feature=+simd128"]
Then I wrote the int8 dot product with explicit core::arch::wasm32 intrinsics — 16 int8 pairs per step, widened to i16 products, pair-added into i32 accumulators:
#[cfg(target_arch = "wasm32")]
#[inline]
fn dot_i8(a: &[i8], b: &[i8]) -> i32 {
use core::arch::wasm32::*;
let mut acc = i32x4_splat(0);
let mut i = 0;
unsafe {
while i < a.len() {
let va = v128_load(a.as_ptr().add(i) as *const v128);
let vb = v128_load(b.as_ptr().add(i) as *const v128);
let lo = i16x8_extmul_low_i8x16(va, vb); // 8 x i16 products
let hi = i16x8_extmul_high_i8x16(va, vb);
acc = i32x4_add(acc, i32x4_extadd_pairwise_i16x8(lo));
acc = i32x4_add(acc, i32x4_extadd_pairwise_i16x8(hi));
i += 16;
}
}
i32x4_extract_lane::<0>(acc) + i32x4_extract_lane::<1>(acc)
+ i32x4_extract_lane::<2>(acc)
+ i32x4_extract_lane::<3>(acc)
}
3.2 s → ~1.16 s, and eventually ~1 s per request. Usable, not fast. (The second post is all about the next 4×.)
Release profile, tuned for speed:
[profile.release] opt-level = 3 lto = true codegen-units = 1 panic = "abort" strip = true
Step 5: tokenizer, decoding, and a deterministic safety net
Tokenizer
The model uses a SentencePiece Unigram tokenizer. Hugging Face’s tokenizers crate compiles to Wasm if you turn off the default features:
tokenizers = { version = "0.20", default-features = false, features = ["unstable_wasm"] }
and tokenizer.json is embedded the same way as the weights.
From logits to spans
The 89 labels are BIOES tags (B-, I-, E-, S- per category, plus O). The handler takes a softmax per token, decodes BIOES into spans, maps token offsets back to character offsets, and assigns numbered placeholders ([GIVEN_NAME_1], [GIVEN_NAME_2], …).
The classifier cuts numbers short
Testing structured data showed a typical token-classifier weakness: spans cut short at token edges. A 16-digit card tagged as 4111 1111 only; an IP address losing its 192.. For data with checksums, a deterministic check is simply more reliable than a neural network, so I added src/deterministic.rs:
/// Luhn (ISO/IEC 7812-1) check digit.
fn luhn_valid(digits: &[u8]) -> bool {
if digits.len() < 13 {
return false;
}
let mut sum = 0u32;
let mut double = false;
for &c in digits.iter().rev() {
let mut d = (c - b'0') as u32;
if double {
d *= 2;
if d > 9 { d -= 9; }
}
sum += d;
double = !double;
}
sum % 10 == 0
}
/// ISO-13616 IBAN: move the first four characters to the end, map letters to
/// two-digit numbers, and check the whole thing is 1 mod 97.
fn iban_valid(s: &[u8]) -> bool {
if s.len() < 15 || s.len() > 34 { return false; }
let rotated = s[4..].iter().chain(s[..4].iter());
let mut rem: u32 = 0;
for &c in rotated {
let val = if c.is_ascii_digit() { (c - b'0') as u32 }
else if c.is_ascii_uppercase() { (c - b'A') as u32 + 10 }
else { return false };
rem = if val >= 10 { (rem * 100 + val) % 97 } else { (rem * 10 + val) % 97 };
}
rem == 1
}
- Luhn for credit cards and IMEIs
- ISO 13616 mod-97 for IBANs
- Strict dotted-quad validation for IPv4
Matches that pass the checksum override the classifier’s spans, so a valid card number is always redacted in full.
Step 6: packaging for Spin
A route gotcha
Our first manifest used route = "/". POST /redact returned 404: "/" matches only the root path. The wildcard form is needed:
spin_manifest_version = 2 [application] name = "redact-wasm" version = "0.1.0" [[trigger.http]] route = "/..." component = "redact-wasm" [component.redact-wasm] source = "target/wasm32-wasip1/release/redact_wasm.wasm" allowed_outbound_hosts = [] [component.redact-wasm.build] command = "cargo build --target wasm32-wasip1 --release"
allowed_outbound_hosts = [] is a nice property: the component cannot send the text it’s redacting anywhere.
The handler
#[http_component]
fn handle(req: Request) -> anyhow::Result<impl IntoResponse> {
if req.method() != &Method::Post {
return Ok(Response::builder().status(405).build());
}
// 1. parse {"text": "..."}
// 2. tokenize
// 3. model.forward() per 256-token window
// 4. BIOES decode + deterministic overrides
// 5. return JSON
}
Build, run, deploy
cd redact-wasm
python3 extract.py # once: downloads redact.tflite, writes assets/model.bin
spin build
spin up # http://localhost:3000
curl -s -X POST http://localhost:3000/redact \
-H "Content-Type: application/json" \
-d '{"text": "My name is John Doe, I live in Paris, and my email is john.doe@example.com."}'
spin aka deploy # to Akamai Functions
Two more small gotchas I hit:
- Deploy from the right directory. Running
spin aka deployfrom the JS project deployed the “Hello from Spin!” stub instead of the model. - Port 3000 stays busy if an old
spin upis still running:kill -9 $(lsof -t -i :3000).
The deployed component returned the same redaction as locally, running the full 23M-parameter forward pass inside a ~28 MB Wasm component on Akamai.
What I’d tell someone porting their own model
- Know the platform before choosing tools. No nested Wasm in JS, no FFI, and a new instance per request rule out most “just use the SDK” approaches.
- Treat start-up as per-request cost. Flat weights +
include_bytes!+ slices = no load time. - Get a golden reference first. Every one of our surprises (mask, padding, quantisation, folded LayerNorm) would silently cost accuracy without a LiteRT comparison.
- Emulate the reference’s arithmetic, not the ideal math. Hybrid int8 quantisation is part of the model.
- Trust the graph, not conventions. The attention mask and pad token did not follow the Hugging Face defaults.
- Assert everything in the extractor: op counts, quantisation scheme, total byte size.
- Add deterministic checks for structured PII. Checksums beat classifiers on card numbers and IBANs.
What fits on Akamai Functions?
| Good fit | Poor fit |
|---|---|
| 5–40M parameter encoders (MiniLM, MobileBERT, TinyBERT, DistilBERT-size) | Multi-billion parameter LLMs |
| 10–40 MB components | Anything needing a GPU |
| Classification, NER, embeddings, small vision backbones | Models needing native runtimes or dynamic graph loading |
At this point I had a correct, deployed model taking about 1 second per request. The next post covers how I got it to ~260 ms without changing a single bit of its output.
Written By

I’m an Enterprise Architect at Akamai Technologies with 15+ years of experience across mobile engineering, edge infrastructure, security, and AI systems. Having launched 45+ apps on the App Store and Play Store (iOS, Android, Flutter, React Native), I specialize in mobile SDK internals, Frida-based security, and high-concurrency edge runtimes like Akamai EdgeWorkers, Fermyon, and HarperDB. In the AI space, I focus on Agentic AI frameworks (LangGraph, MCP), WASM-based Edge AI guardrails, self-hosted LLM inference, and real-time voice pipelines.