Replacing Moderation APIs With In-House Text Classifiers: Silver Labels, Shadow Scoring and a Regression Gate
How I trained ModernBERT and XLM-R classifiers on silver labels from third-party APIs, ran them in shadow next to the incumbent, and built a baseline regression gate for the model registry.
For several months I worked on text classifiers for a harmful-content detection system at a nonprofit. Posts came in, got scored, and the ones above a threshold were flagged for human review. When I started, most of the scoring came from two third-party moderation APIs. The goal was to replace them with models we trained and served ourselves.
This post covers how the training data was built, how I checked that the new models agreed with the old scores before switching anything over, and the retraining pipeline with a regression gate that I added later.
Silver labels
We had no large hand-labeled set. We did have months of production text that the third-party APIs had already scored. Those scores became silver labels: noisy targets from a teacher we were trying to replace.
I trained two families. ModernBERT-large handled English text, with one binary classifier and one multilabel regression model that predicted the APIs' score dimensions directly. XLM-RoBERTa-large handled everything else, with a language detector routing each text to one model or the other.
The first runs were on single-GPU EC2 instances, driven by a shell script that launched the box, uploaded data, installed dependencies and started training. ModernBERT used a learning rate of 2e-5, XLM-R used 1e-5, and both used a 512-token max length. The binary trainers shared one core:
n_pos = train_df["label"].sum()
n_neg = len(train_df) - n_pos
class_weights = torch.tensor(
[len(train_df) / (2 * n_neg), len(train_df) / (2 * n_pos)], dtype=torch.float32
)
args = TrainingArguments(
learning_rate=lr,
gradient_accumulation_steps=4,
lr_scheduler_type="cosine",
warmup_ratio=0.1,
weight_decay=0.01,
fp16=True,
gradient_checkpointing=True,
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="pr_auc",
)
class WeightedTrainer(Trainer):
def compute_loss(self, model, inputs, return_outputs=False, **kwargs):
labels = inputs.pop("labels")
outputs = model(**inputs)
loss = torch.nn.CrossEntropyLoss(weight=class_weights.to(outputs.logits.device))(
outputs.logits, labels
)
return (loss, outputs) if return_outputs else loss
Positives were rare, so the loss was class-weighted, and checkpoints were chosen on PR-AUC instead of accuracy or F1. On the English binary classifier, my first ModernBERT run reached F1 0.759 (precision 0.673, recall 0.868, PR-AUC 0.720). A second run a day later reached F1 0.838 (precision 0.816, recall 0.861, PR-AUC 0.907). The eval split was not the same size in the two runs, so the jump is not a clean comparison. The second-run numbers became the recorded baseline for that model.
One preprocessing detail mattered for the multilingual model. The English pipeline decoded leet speak and segmented hashtags, among other steps, and those implementations were written for English. I removed them from the XLM-R trainer because I expected them to hurt cross-lingual transfer.
Silver-label metrics flatter you
Two smaller binary classifiers trained on silver labels scored F1 0.995 and 0.994 on their held-out splits. My reading was that they matched the teacher very well on text from the same distribution, which says nothing about whether the teacher was right. The comparison notes I wrote about the incumbent model made a related point: a held-out split of the same teacher-labeled data says little about accuracy against human judgment or on different data.
So I also evaluated on public labeled datasets that the models had never seen. At the default 0.5 threshold, the main binary classifier missed about 36% of positives, and the notebook's threshold sweep pointed to about 0.25. On one non-English set, recall at 0.5 was 13% even though AUC-ROC was 0.866. The model ranked those texts well, but its scores sat below 0.5. So I made thresholds configurable per classifier, with 0.25 as the default for the main binary classifier and its multilingual counterpart, instead of hardcoding 0.5.
Exporting training data without slow joins
The silver labels lived in a large, partitioned Postgres table of scores that joined to a content table. The first export was a local script run over an SSH tunnel. I moved it to a Lambda plus Step Functions pipeline that writes CSVs to S3.
The first version joined scores to content directly, and it took 171 seconds to return a small test batch. The version that worked split the query in two:
# 1. IDs and scores only, served by btree indexes on the scores table (under 2 s)
ids = fetch_ids_and_scores(conn, since=window_start, limit=target_rows)
# 2. Fetch the text in batches with an IN clause
for batch in chunks(ids, 5000):
rows = conn.execute(
"SELECT content_id, body FROM content WHERE content_id IN %s",
(tuple(r.content_id for r in batch),),
)
write_rows(merge(batch, rows))
Shadow scoring before cutover
The in-house models did not replace anything at first. Next I shipped a shadow pipeline: every classified item also got in-house scores, stored in a separate JSONB column together with the detected language. This let us compare old and new scores on real traffic without changing what reviewers saw.
Sampling for the batch parity evaluation took several attempts. ORDER BY RANDOM() was too slow, TABLESAMPLE errored on our partitioned table, and sorting by an md5 hash still did a full scan. A streaming cursor with filtering in Python was what worked. Long inputs also made some chunks time out. The fixes were a 120-second timeout per chunk, sorting inputs by length before batching, and dropping the chunk size from 50 to 32. An early sample showed an average Pearson correlation of 0.948 against one API's dimensions and 0.789 against the other's, with 99.2% agreement at the 0.5 threshold. A few weeks later, on a much larger paired sample, Pearson was between 0.962 and 0.975 across the dimensions in that report, and agreement at 0.5 was 97.7% to 99.3%. I read through a sample of the disagreements. In most of them the in-house model scored high and the incumbent scored low, and they were mostly short, informal texts that the in-house model had flagged wrongly.
The shadow data also surfaced two bugs. When the language-detection call failed, the client defaulted to 'en', so non-English text went only to the English models and was stored with the wrong language. I changed it to return null on failure and routed unknown-language items through both model families. Then I nulled the shadow scores on the affected records so the retry workflow would classify them again. The second bug showed up during an outage of the classifier service. The service returned error-only payloads, and the batch writer saved them over good scores. I added a guard so the writer only overwrites existing scores when the new payload contains actual model output.
A regression gate in front of the registry
Retraining ran through the same Step Functions and SageMaker pipeline as the image model. A finished job registered its model as PendingManualApproval, and an engineer approved it. I later added a gate in front of that, in four pieces.
First, a comparison module reads eval_results.json, compares it against baselines.yaml with a default absolute tolerance of 0.01, and exits non-zero on a regression:
def compare(results: dict, baseline: dict, tolerance: float = 0.01) -> list[str]:
regressions = []
for metric, base in baseline["metrics"].items():
got = results.get(metric)
if got is None or got < base - tolerance:
regressions.append(f"{metric}: {got} vs baseline {base}")
return regressions
For multilabel regression the metrics are mean Pearson, MAE and RMSE, and for MAE and RMSE lower is better, so the real module handles both directions.
Second, the training entrypoint runs the comparison and writes the report. A failed comparison does not crash training. The report also goes to its own S3 key, because the state machine cannot read files inside model.tar.gz.
Third, there is a small golden set per model. It is manually verified and immutable, and a new version means a new directory. Its eval output has the same shape as the training-time results, so it can go through the same comparison.
Fourth, the state machine loads the report. On a regression it skips registration and publishes to a separate SNS topic for model-quality alerts. All of this sits behind a flag that is off by default, to stay off until the baselines had been checked against a few real runs. The gate fails open: if loading the report fails, the execution registers the model anyway, and a human still has to approve it before anything deploys. I preferred that over blocking on a plumbing failure. I have no record of the flag being turned on in production before I left.
Things that broke on the way to production
Approved models that never deployed. Approving a model wrote the new artifact path to SSM Parameter Store and rolled the API service. However, the hourly GPU batch job that actually did the multilingual scoring had its model path baked into its AWS Batch job definition at deploy time. Nothing touched that job definition after an approval, so it kept loading the previous weights. I changed the batch job to read the active path from SSM at startup, in the order SSM, then env, then default.
pandas 3. After the training image picked up pandas 3 and a new pyarrow through unpinned dependencies, string label columns came back as a pandas string dtype instead of object. That broke the labels.dtype == object check, and the trainer fell into the numeric branch and crashed. The fix was pd.api.types.is_numeric_dtype, plus pins until CI covered newer versions.
ONNX. I tried to serve INT8 ONNX models. The legacy TorchScript exporter failed on the attention-mask helpers in the current transformers release. The dynamo exporter produced a graph, but quantization then failed with a shape-inference error on the classification head. I left the ONNX path in the code but switched it off, so PyTorch stayed the default in every stage.