How I got a multilingual PII redaction model running on Akamai Functions (Fermyon Spin + WebAssembly), after the obvious approaches failed one by one.

The goal

I wanted an HTTP endpoint that takes text and returns it with personal data replaced by placeholders:

curl -s -X POST https://<app>.fwf.app/redact \
     -H "Content-Type: application/json" \
     -d '{"text": "My name is John Doe, I live in Paris, and my email is john.doe@example.com."}'
{
  "redactedText": "My name is [GIVEN_NAME_1] [SURNAME_1], I live in [CITY_1], and my email is [EMAIL_1].",
  "items": [
    { "label": "GIVEN_NAME", "original": "John",  "placeholder": "[GIVEN_NAME_1]", "start": 11, "end": 15 },
    { "label": "SURNAME",    "original": "Doe",   "placeholder": "[SURNAME_1]",    "start": 16, "end": 19 },
    { "label": "CITY",       "original": "Paris", "placeholder": "[CITY_1]",       "start": 31, "end": 36 },
    { "label": "EMAIL",      "original": "john.doe@example.com", "placeholder": "[EMAIL_1]", "start": 54, "end": 74 }
  ]
}

The model is Desert Ant Labs’ redact: a multilingual PII token classifier.

PropertyValue
ArchitectureBERT-style encoder (XLM-RoBERTa / MiniLM lineage)
Parameters~23M, int8 weights
Layers / hidden / heads / FFN6 / 384 / 12 / 1536
Sequence lengthfixed 256 tokens
Output89 BIOES labels over 20 PII categories (GIVEN_NAME, SURNAME, CITY, EMAIL, CREDIT_CARD, IBAN, …)
Distributionnpm package @desert-ant-labs/redact and redact.tflite on Hugging Face

And I wanted it on Akamai Functions: WebAssembly components running on Spin, close to users, with no containers, no GPUs and no sidecars.

This post covers how I got there, including the dead ends. The follow-up post covers how I then took latency from ~1 s down to ~260 ms.

Understanding the platform first

Three properties of Akamai Functions / Spin decide almost everything below:

  1. Code runs as a WebAssembly component in a WASI sandbox. No native libraries, no dlopen, no C FFI to the host.
  2. JavaScript runs inside StarlingMonkey (SpiderMonkey compiled to Wasm, via ComponentizeJS). There is no DOM, no Web Workers, and — the one that bit us — no nested WebAssembly API.
  3. A fresh instance is created for every request. Whatever you set up at start-up, you pay for on every request.

I didn’t fully understand all three at the start. I learned them in order.

Attempt 1: “Just use the npm package” (JavaScript)

The model ships as an npm package, and Spin has a JavaScript template. The obvious first try:

import { Redact } from "@desert-ant-labs/redact";

let model = null;

export async function handleRequest(request) {
    if (!model) model = await Redact.load();
    const { text } = await request.json();
    const result = await model.redaction(text);
    return new Response(JSON.stringify(result), {
        headers: { "content-type": "application/json" },
    });
}

Problem 1: dynamic Node imports

Componentising (j2w) failed straight away. The package’s default (Node) entry point did:

const fs = await import("node:fs");

ComponentizeJS rejects that:

Dynamic module import is disabled or not supported in this context

The Node build also loads the model through koffi, a native FFI library that calls into .dylib/.so files. That can never work inside a Wasm sandbox.

Problem 2: the browser build hangs silently

So I pointed the bundler at the browser build (conditions: ['browser']). It compiled. It deployed. Every request failed with:

ERROR: no fetch-event handler triggered

No stack trace, no exception. Reading @desert-ant-labs/redact/browser.js explained it. At the top level of the module:

const sdk = await createWasmSdk(/* ... */);   // top-level await

createWasmSdk uses LiteRT.js, which loads its own Wasm runtime by injecting <script> tags and compiling Wasm with the WebAssembly API. Inside StarlingMonkey:

typeof WebAssembly === "undefined"   // true

So the top-level promise never resolved, the module never finished initialising, and addEventListener('fetch', ...) was never registered. Spin saw a component with no handler.

Lesson: JavaScript ML packages that depend on nested WebAssembly, DOM script injection, Web Workers or native FFI cannot run in Spin’s JS runtime. The JS project (redact-akamai-function/) was left as a stub, and I moved to Rust.

Attempt 2: an off-the-shelf Rust runtime (tract-tflite)

Rust compiles straight to wasm32-wasip1, and tract is a pure-Rust inference engine with a TFLite frontend. I downloaded redact.tflite and tried:

let model = tract_tflite::tflite()
    .model_for_path("redact.tflite")?
    .into_optimized()?
    .into_runnable()?;

It failed on load:

Error: Translating proto model to model
Caused by: Error in TF file for operator Operator { ... }. No prior computation nor constant for input 115

Tensor 115 was the word embedding table. Why would a constant be “missing”?

Digging into the flatbuffer

I inspected the file with the Python tflite package:

import tflite

buf = open("redact.tflite", "rb").read()
model = tflite.Model.GetRootAsModel(buf, 0)
sub = model.Subgraphs(0)

for i in range(sub.TensorsLength()):
    t = sub.Tensors(i)
    b = model.Buffers(t.Buffer())
    print(i, t.Name().decode(), t.Type(), t.ShapeAsNumpy(),
          "inline:", b.DataLength(), "offset:", b.Offset(), "size:", b.Size())

101 of the 122 buffers had DataLength() == 0. Their data wasn’t inside the flatbuffer at all. They used TFLite’s external buffer offset extension: Buffer.offset and Buffer.size point to raw bytes appended after the end of the flatbuffer. That’s how large models get past flatbuffers’ 2 GB limit, and some exporters use it for all big tensors. tract-tflite only reads inline buffers, so to it every weight was missing.

Even if tract had loaded the model, there was a second, deeper problem: Spin creates a new instance per request. Any runtime that parses a model graph, allocates tensors and optimises the plan at start-up does that on every request. For graph-based runtimes (tract, ONNX Runtime, the TFLite interpreter compiled to Wasm) that is typically hundreds of milliseconds to seconds before any real work starts.

Lesson: On a per-request-instance platform, start-up cost is per-request cost. I needed a format that needs no parsing at all.

The design that worked: flat weights + hand-written forward pass

The plan:

  1. Extract every weight from the .tflite into a single flat binary file in a fixed, known order.
  2. Embed that file into the Wasm binary with include_bytes!, so the weights are in memory as soon as the instance starts.
  3. Write the forward pass by hand in Rust — it’s only six BERT layers.
  4. Verify against the official LiteRT runtime until the outputs match.
redact.tflite ──extract.py──▶ assets/model.bin ──include_bytes!──▶ Wasm component
                                                                        │
                                                Model::load (slices, <1 ms)
                                                                        │
                                              forward(): 6 layers in Rust + Wasm SIMD

Step 1: reverse-engineering the graph (extract.py)

There’s no metadata that says “this is the Q projection of layer 3”. I had to recover the architecture from the graph’s operators.

Reading external buffers

First, a raw() helper that understands both inline and external buffers:

def raw(idx):
    """Bytes of a constant tensor, honouring the external-buffer extension."""
    t = sub.Tensors(idx)
    b = model.Buffers(t.Buffer())
    off = b.Offset()
    if off and off > 1:
        return buf[off:off + b.Size()]        # data lives past the flatbuffer
    assert b.DataLength() > 0, f"tensor {idx} is not constant"
    return b.DataAsNumpy().tobytes()

Finding the matrix multiplies

A BERT layer has six dense layers: Q, K, V, output projection, FFN-in, FFN-out. With 6 layers plus the classifier I expected 37 FULLY_CONNECTED ops, and found exactly 37. In topological order they come out as [Q, K, V, O, FF1, FF2] × 6 + classifier:

fcs = []
for k in range(sub.OperatorsLength()):
    op = sub.Operators(k)
    if opnames[op.OpcodeIndex()] == "FULLY_CONNECTED":
        ins = [int(x) for x in op.InputsAsNumpy()]
        fcs.append((ins[1], ins[2]))        # (weights, bias)
assert len(fcs) == 37, len(fcs)

Checking the quantisation scheme

Each weight tensor carries its own quantisation info. I asserted the assumptions rather than hoping:

def scales(idx):
    q = sub.Tensors(idx).Quantization()
    s = np.array([q.Scale(i) for i in range(q.ScaleLength())], dtype=np.float32)
    assert q.QuantizedDimension() == 0, "expected per-output-channel quantisation"
    zp = np.array([q.ZeroPoint(i) for i in range(q.ZeroPointLength())])
    assert (zp == 0).all(), "expected symmetric quantisation"
    return s

So: int8 weights, one float scale per output row, zero-point 0.

Finding LayerNorm

TFLite had no LayerNorm op here; it was decomposed into mean / subtract / variance / rsqrt / MUL / ADD. The affine part (γ, β) shows up as a MUL by a constant [384] vector followed by an ADD of another. I found 12 such pairs: 1 after the embeddings, and 2 per layer except the very last one.

Folded embeddings

Position embeddings and token-type embeddings were two constant [1, 256, 384] tensors added after the word-embedding lookup. Since the sequence length is fixed and token type is always 0, I folded them into one bias:

put_f32(f32(bias_consts[0]).reshape(256, 384) + f32(bias_consts[1]).reshape(256, 384))

The final layout

word_emb      int8 [31475*384]   + word_scale  f32 [31475]
emb_bias      f32  [256*384]                      (position + token_type, folded)
emb_ln_gamma  f32  [384]         + emb_ln_beta f32 [384]
6 x layer:
    q,k,v,o : w int8[384*384] + scale f32[384]  + bias f32[384]
    attn_ln : gamma f32[384]   + beta f32[384]
    ffn1    : w int8[1536*384] + scale f32[1536] + bias f32[1536]
    ffn2    : w int8[384*1536] + scale f32[384]  + bias f32[384]
    ffn_ln  : gamma f32[384]   + beta f32[384]
cls           int8 [89*384]      + scale f32[89] + bias f32[89]

The script computes the expected size and checks it — 23,463,060 bytes. That check caught an off-by-one early on:

print(f"wrote {DST}: {len(out):,} bytes (expected {expected:,}) match={len(out)==expected}")

Step 2: loading without parsing

On the Rust side, “loading” is just walking a cursor over the embedded bytes. Weight matrices are slices into the binary, not copies:

static MODEL_BLOB: &[u8] = include_bytes!("../assets/model.bin");

struct Cursor<'a> { b: &'a [u8], at: usize }

impl<'a> Cursor<'a> {
    fn i8s(&mut self, n: usize) -> &'a [i8] {
        let s = &self.b[self.at..self.at + n];
        self.at += n;
        // i8 and u8 have identical size and alignment.
        unsafe { core::slice::from_raw_parts(s.as_ptr() as *const i8, n) }
    }

    fn qmat(&mut self, out: usize, inp: usize) -> QMat<'a> {
        let w = self.i8s(out * inp);
        let scale = self.f32s(out);
        let bias = self.f32s(out);
        QMat { w, scale, bias, out, inp }
    }
}

impl<'a> Model<'a> {
    pub fn load(blob: &'a [u8]) -> Model<'a> {
        let mut c = Cursor { b: blob, at: 0 };
        let emb = c.i8s(VOCAB * D);
        // ... embeddings, 6 layers, classifier ...
        assert_eq!(c.at, blob.len(), "model.bin size mismatch");
        Model { /* ... */ }
    }
}

Model::load takes under a millisecond. The 22 MB of weights cost nothing at start-up because they are already part of the component’s memory image.

Step 3: a golden reference, and three surprises

Before trusting any Rust output, I set up a reference: the official LiteRT Python runtime (ai-edge-litert) running redact.tflite on the benchmark sentence, saving input ids and logits.

np.save("golden_logits.npy", reference_logits)
json.dump({"ids": input_ids, "text": text}, open("golden_input.json", "w"))

Then I compared our Rust forward pass against it. Three surprises came out of that.

Surprise A: the attention mask must be all ones

I started with the standard Hugging Face convention: attention_mask = 1 for real tokens, 0 for padding. The model found the email address — and nothing else. No John, no Doe, no Paris.

Setting the mask to 1 everywhere, including padding, restored every detection and matched LiteRT. The exported graph was traced with an all-ones mask, and that’s the behaviour it was calibrated on.

Surprise B: padding is load-bearing

Since the mask is all ones, the model attends to padding tokens. So the padding itself affects the output:

  • Pad token 0 instead of 1 → logits shifted.
  • Trimming the sequence below 256 → logits shifted by up to 13.28.

So every request runs a full 256-token window padded with id 1:

pub const PAD_ID: u32 = 1;

// The graph is fixed at 256 tokens and attends over the whole window, so
// longer input is processed as successive full windows.
let mut window = vec![PAD_ID; SEQ];
window[..chunk_end - chunk_start].copy_from_slice(&ids[chunk_start..chunk_end]);
let logits = model.forward(&window, chunk_end - chunk_start);

Surprise C: float is “too accurate”

Our first matrix multiply dequantised the int8 weights and multiplied in f32. Mathematically that’s more precise than int8 — and it agreed with LiteRT on only 86% of token labels. The error grew layer by layer.

The reason is TFLite’s hybrid quantisation. For an int8-weight FULLY_CONNECTED with float input, TFLite:

  1. Quantises each input row to int8 on the fly: scale = max|x| / 127, q = round(x / scale).
  2. Does an integer dot product with the int8 weights.
  3. Dequantises: y = bias + acc × scale_in × scale_w.

The model was calibrated with that rounding in the loop. Reproducing it exactly:

/// TFLite hybrid FullyConnected.
///
/// Emulating the dynamic activation quantisation matters: a plain f32
/// matmul is *more precise* than the reference and drifts from it by about
/// one quantisation step per layer.
fn quantize_row(xr: &[f32], q: &mut [i8]) -> f32 {
    let mut amax = 0f32;
    for &a in xr {
        let m = if a < 0.0 { -a } else { a };
        if m > amax { amax = m; }
    }
    if amax == 0.0 { q.fill(0); return 0.0; }
    let inv = 127.0 / amax;
    for k in 0..xr.len() {
        let v = (xr[k] * inv).round();
        q[k] = if v > 127.0 { 127 } else if v < -127.0 { -127 } else { v as i8 };
    }
    amax / 127.0
}

// ... then per output o:
y[o] = bias[o] + (acc as f32) * scale_in * scale_w[o];

Result: 100% label agreement with LiteRT, max logit difference 0.31.

Bonus surprise: folded LayerNorm affine

Tracing the graph more closely showed that the LayerNorm γ/β were not applied before Q/K/V and the FFN. They had been folded into those layers’ weights. The graph had two versions of each normalised tensor:

  • Pre-affine (just normalised) → feeds Q/K/V or the FFN.
  • Post-affine (× γ + β) → feeds the residual connection.

So the layer structure is:

layer_norm(&x, n, &mut norm);                              // pre-affine
apply_affine(&norm, n, &ly.attn_gamma, &ly.attn_beta, &mut affine); // post-affine

ly.ff1.matmul(&norm, ...);          // FFN consumes the pre-affine tensor
// ...
x[i] = proj[i] + affine[i];         // residual adds the post-affine tensor

The final layer’s LayerNorm has an identity affine (γ=1, β=0), which extract.py writes explicitly.

Step 4: making it fast enough to be usable

The first correct version was scalar Rust, and a request took ~3.2 s. I tried relying on LLVM’s auto-vectoriser with SIMD enabled; the int8 dot product barely improved.

Enabling Wasm SIMD is one line in .cargo/config.toml:

[target.wasm32-wasip1]
rustflags = ["-C", "target-feature=+simd128"]

Then I wrote the int8 dot product with explicit core::arch::wasm32 intrinsics — 16 int8 pairs per step, widened to i16 products, pair-added into i32 accumulators:

#[cfg(target_arch = "wasm32")]
#[inline]
fn dot_i8(a: &[i8], b: &[i8]) -> i32 {
    use core::arch::wasm32::*;
    let mut acc = i32x4_splat(0);
    let mut i = 0;
    unsafe {
        while i < a.len() {
            let va = v128_load(a.as_ptr().add(i) as *const v128);
            let vb = v128_load(b.as_ptr().add(i) as *const v128);
            let lo = i16x8_extmul_low_i8x16(va, vb);   // 8 x i16 products
            let hi = i16x8_extmul_high_i8x16(va, vb);
            acc = i32x4_add(acc, i32x4_extadd_pairwise_i16x8(lo));
            acc = i32x4_add(acc, i32x4_extadd_pairwise_i16x8(hi));
            i += 16;
        }
    }
    i32x4_extract_lane::<0>(acc) + i32x4_extract_lane::<1>(acc)
        + i32x4_extract_lane::<2>(acc)
        + i32x4_extract_lane::<3>(acc)
}

3.2 s → ~1.16 s, and eventually ~1 s per request. Usable, not fast. (The second post is all about the next 4×.)

Release profile, tuned for speed:

[profile.release]
opt-level = 3
lto = true
codegen-units = 1
panic = "abort"
strip = true

Step 5: tokenizer, decoding, and a deterministic safety net

Tokenizer

The model uses a SentencePiece Unigram tokenizer. Hugging Face’s tokenizers crate compiles to Wasm if you turn off the default features:

tokenizers = { version = "0.20", default-features = false, features = ["unstable_wasm"] }

and tokenizer.json is embedded the same way as the weights.

From logits to spans

The 89 labels are BIOES tags (B-, I-, E-, S- per category, plus O). The handler takes a softmax per token, decodes BIOES into spans, maps token offsets back to character offsets, and assigns numbered placeholders ([GIVEN_NAME_1], [GIVEN_NAME_2], …).

The classifier cuts numbers short

Testing structured data showed a typical token-classifier weakness: spans cut short at token edges. A 16-digit card tagged as 4111 1111 only; an IP address losing its 192.. For data with checksums, a deterministic check is simply more reliable than a neural network, so I added src/deterministic.rs:

/// Luhn (ISO/IEC 7812-1) check digit.
fn luhn_valid(digits: &[u8]) -> bool {
    if digits.len() < 13 {
        return false;
    }
    let mut sum = 0u32;
    let mut double = false;
    for &c in digits.iter().rev() {
        let mut d = (c - b'0') as u32;
        if double {
            d *= 2;
            if d > 9 { d -= 9; }
        }
        sum += d;
        double = !double;
    }
    sum % 10 == 0
}

/// ISO-13616 IBAN: move the first four characters to the end, map letters to
/// two-digit numbers, and check the whole thing is 1 mod 97.
fn iban_valid(s: &[u8]) -> bool {
    if s.len() < 15 || s.len() > 34 { return false; }
    let rotated = s[4..].iter().chain(s[..4].iter());
    let mut rem: u32 = 0;
    for &c in rotated {
        let val = if c.is_ascii_digit() { (c - b'0') as u32 }
                  else if c.is_ascii_uppercase() { (c - b'A') as u32 + 10 }
                  else { return false };
        rem = if val >= 10 { (rem * 100 + val) % 97 } else { (rem * 10 + val) % 97 };
    }
    rem == 1
}
  • Luhn for credit cards and IMEIs
  • ISO 13616 mod-97 for IBANs
  • Strict dotted-quad validation for IPv4

Matches that pass the checksum override the classifier’s spans, so a valid card number is always redacted in full.

Step 6: packaging for Spin

A route gotcha

Our first manifest used route = "/". POST /redact returned 404: "/" matches only the root path. The wildcard form is needed:

spin_manifest_version = 2

[application]
name = "redact-wasm"
version = "0.1.0"

[[trigger.http]]
route = "/..."
component = "redact-wasm"

[component.redact-wasm]
source = "target/wasm32-wasip1/release/redact_wasm.wasm"
allowed_outbound_hosts = []

[component.redact-wasm.build]
command = "cargo build --target wasm32-wasip1 --release"

allowed_outbound_hosts = [] is a nice property: the component cannot send the text it’s redacting anywhere.

The handler

#[http_component]
fn handle(req: Request) -> anyhow::Result<impl IntoResponse> {
    if req.method() != &Method::Post {
        return Ok(Response::builder().status(405).build());
    }
    // 1. parse {"text": "..."}
    // 2. tokenize
    // 3. model.forward() per 256-token window
    // 4. BIOES decode + deterministic overrides
    // 5. return JSON
}

Build, run, deploy

cd redact-wasm
python3 extract.py          # once: downloads redact.tflite, writes assets/model.bin
spin build
spin up                     # http://localhost:3000

curl -s -X POST http://localhost:3000/redact \
     -H "Content-Type: application/json" \
     -d '{"text": "My name is John Doe, I live in Paris, and my email is john.doe@example.com."}'

spin aka deploy             # to Akamai Functions

Two more small gotchas I hit:

  • Deploy from the right directory. Running spin aka deploy from the JS project deployed the “Hello from Spin!” stub instead of the model.
  • Port 3000 stays busy if an old spin up is still running: kill -9 $(lsof -t -i :3000).

The deployed component returned the same redaction as locally, running the full 23M-parameter forward pass inside a ~28 MB Wasm component on Akamai.

What I’d tell someone porting their own model

  1. Know the platform before choosing tools. No nested Wasm in JS, no FFI, and a new instance per request rule out most “just use the SDK” approaches.
  2. Treat start-up as per-request cost. Flat weights + include_bytes! + slices = no load time.
  3. Get a golden reference first. Every one of our surprises (mask, padding, quantisation, folded LayerNorm) would silently cost accuracy without a LiteRT comparison.
  4. Emulate the reference’s arithmetic, not the ideal math. Hybrid int8 quantisation is part of the model.
  5. Trust the graph, not conventions. The attention mask and pad token did not follow the Hugging Face defaults.
  6. Assert everything in the extractor: op counts, quantisation scheme, total byte size.
  7. Add deterministic checks for structured PII. Checksums beat classifiers on card numbers and IBANs.

What fits on Akamai Functions?

Good fitPoor fit
5–40M parameter encoders (MiniLM, MobileBERT, TinyBERT, DistilBERT-size)Multi-billion parameter LLMs
10–40 MB componentsAnything needing a GPU
Classification, NER, embeddings, small vision backbonesModels needing native runtimes or dynamic graph loading

At this point I had a correct, deployed model taking about 1 second per request. The next post covers how I got it to ~260 ms without changing a single bit of its output.


Written By
Fareeth John

I’m an Enterprise Architect at Akamai Technologies with 15+ years of experience across mobile engineering, edge infrastructure, security, and AI systems. Having launched 45+ apps on the App Store and Play Store (iOS, Android, Flutter, React Native), I specialize in mobile SDK internals, Frida-based security, and high-concurrency edge runtimes like Akamai EdgeWorkers, Fermyon, and HarperDB. In the AI space, I focus on Agentic AI frameworks (LangGraph, MCP), WASM-based Edge AI guardrails, self-hosted LLM inference, and real-time voice pipelines.

Leave a Reply

Your email address will not be published. Required fields are marked *