Laya Multilingual for Stellar operations

A fine-tune of convaiinnovations/laya-multilingual (322M parameters) at revision 1720e3e3. It is tuned to read the one-line sentences the Stellar radar writes about live Stellar network operations: payments, offers, path payments, trustlines, account creation, contract calls. It puts each operation into a lane and answers yes/no questions that viewers ask about it.

On 1,500 benchmark operations, lane agreement with the ledger rises from 93.5% to 99.8%. Question AUC rises from 0.910 to 0.988. On 2,405 test operations whose sentences never appeared in training, every held-out set improves.

Usage

LAYAD_MODEL=goktugoguz/laya-multilingual-stellar-mlx layad serve

This is the fp16 MLX conversion of goktugoguz/laya-multilingual-stellar; on the validation rows the two give the same lane every time.

What it reads

One English sentence per operation, rendered by the radar, for example: an account swaps 67.03 1x for 0.096 XLM through the order book, or payment of 7.00 XLM from an account to an account, or an account places a buy offer for 100.00 XLM, paying in USDC.

The radar's older templates (for example path payment: an account sends ... and an account receives ...) are not what it was trained on.

It was trained on the radar's corrected sentences. In those, a buy offer states the amount it buys ("places a buy offer for 100.00 XLM, paying in USDC"), and an issued token coded XLM is written XLM-token. Other text is out of its domain.

  • Lane: What kind of economic activity is this Stellar operation?, with these options verbatim and in this order:
  1. remittance: one account pays another account for something
  2. payroll: a wage or salary paid on a schedule
  3. trade: an offer on the order book, a path payment, or a swap between two assets
  4. contract: a smart contract call
  5. onboarding: setting up an account: creating one, or opening or removing a trustline for an asset
  6. spam: a transfer of zero or near-zero value that nobody asked for
  7. other: a change to an account's own settings, or anything else
  • Viewer questions are noul questions asked as written.

Benchmark

The benchmark is 1,500 live operations whose lanes and answers were computed from the ledger's raw fields. A hand audit checked 60 of 60 labels, and 20 operations were cross-checked against stellar.expert.

Base Fine-tune
Lanes agreeing with the ledger (1,255 it settles) 93.5% 99.8%
Unsettled non-zero payments filed as remittance, payroll or spam (244) 78.3% 92.2%
Viewer questions, mean AUC (18) 0.910 0.988
Viewer questions, balanced accuracy at 0.5 (the line the page uses) 0.781 0.933
Stuck rows 1 0

Held out, on 2,405 test operations whose sentences never appeared in training (mean AUC):

Set n Base Fine-tune
Held-out wordings of taught topics 32 0.891 0.996
Topics never taught 15 0.828 0.918
Spanish and German (never trained) 8 0.886 0.999
Russian (never trained) 4 0.809 0.990

No held-out topic wording falls more than 0.05 below the base.

The radar's question gate

Before answering a new question, the radar asks it of fixed probe operations and turns it away if the answers barely differ. With the probe sentences the radar uses today, which are written in an older template this model never saw, the gate refuses 5 of 34 answerable questions (base: 1). With probes rendered in the current template, it refuses 2 (base: 2). Both turn away all 14 nonsense questions. Ship this model with the rendered probes. Each extra refusal under the old probes falls into one of two kinds: a question the base passed on noise, whose top answers were unrelated operations, or a topic with a single probe. A single probe caps a sharp model's spread just under the bar.

Training

  • 18,523 labelled questions over the radar's sentences from a training capture of live operations plus synthetic ones. A hand audit checked 60 of 60 labels. Held-out wordings, topics and languages never reached training.
  • Replay: out-of-bank questions trained toward the base model's own answers (learning without forgetting), capped at 30% of the taught questions. Lane replay, for the shapes the ledger cannot settle, was trained toward the base's own lane distribution. Without it, those payments collapsed into "other".
  • Full fine-tune with frozen token embeddings, soft cross-entropy, 3 epochs; 24,301 training items. Trained on an M4 Pro Mac mini under a supervisor that paused whenever the live radar's model slowed or its queue rose.
  • The released weights are 50% fine-tune and 50% base (WiSE-FT).

Limits

  • Measured on Stellar operations as this radar renders them; anything else is unmeasured.
  • Payments the ledger cannot settle (remittance vs payroll) are a guess for any model. This one files more of them as spam than the base did (182 vs 148 of 244). It also no longer files them as onboarding, a known base error (32).
  • "Is a million or more of anything moving?" stays at balanced accuracy 0.50 at the 0.5 line (AUC ~0.9).
  • English only; Turkish was not a target.

License and attribution

Apache-2.0, as is the base model. Built on convaiinnovations/laya-multilingual by Convai Innovations. Converted with laya-mlx. See NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for goktugoguz/laya-multilingual-stellar-mlx

Finetuned
(70)
this model

Dataset used to train goktugoguz/laya-multilingual-stellar-mlx