Turkish Contact Center Intent Classifier

A Turkish BERT model fine-tuned for intent classification in contact-center and customer-service conversations.

The model is based on dbmdz/bert-base-turkish-cased and classifies Turkish customer utterances into 10 intent categories.

Disclaimer

This is a personal/academic project. It is not affiliated with, endorsed by, or related to any company or institution, and it does not represent the data, systems, or practices of any real contact center or telecom operator.

The model was trained only on synthetic data. No real customer conversations, personal data, or proprietary datasets were used.

Model Details

  • Model type: BERT for sequence classification
  • Base model: dbmdz/bert-base-turkish-cased
  • Language: Turkish
  • Task: Intent classification
  • Number of labels: 10
  • Maximum sequence length used during training: 128
  • Release: v0.1

Supported Intents

ID Intent
0 cihaz_sorunu
1 fatura_itirazi
2 genel_bilgi
3 hat_aktivasyonu
4 internet_arizasi
5 iptal_talebi
6 numara_tasima
7 odeme_sorunu
8 paket_degistirme
9 roaming

Intended Use

The model is intended for experimentation and prototyping of Turkish contact-center intent classification systems.

Example inputs include:

  • internet connectivity problems
  • billing disputes
  • line activation
  • subscription cancellation
  • number porting
  • payment problems
  • package changes
  • roaming questions
  • device problems
  • general service information

Usage

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_name = "yasinunal/turkish-contact-center-intent-classifier"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

model.eval()

text = "İnternetim sürekli kopuyor."

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=128,
)

with torch.no_grad():
    outputs = model(**inputs)

predicted_id = outputs.logits.argmax(dim=-1).item()
predicted_intent = model.config.id2label[predicted_id]

print(predicted_intent)

Expected output:

internet_arizasi

A higher-level Transformers pipeline can also be used:

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="yasinunal/turkish-contact-center-intent-classifier",
)

result = classifier("Faturamda anlamadığım ekstra bir ücret var.")

print(result)

Training Data

The v0.1 model was trained on a Turkish contact-center intent dataset containing 856 training examples.

The training corpus consists of:

  • 800 base training examples
  • 40 contrastive boundary examples
  • 16 targeted repair examples

The additional examples were introduced to strengthen selected intent boundaries.

The dataset contains synthetic and curated examples designed around Turkish contact-center scenarios.

Evaluation

The model was evaluated on held-out in-scope examples and on a separate boundary-oriented benchmark.

Evaluation Result
In-scope test accuracy 100/100 (100%)
Boundary benchmark accuracy 76/80 (95%)

These results should be interpreted in the context of the current datasets. They do not imply 100% accuracy on real-world Turkish contact-center traffic.

Performance may differ with real conversations, ASR transcripts, previously unseen customer goals, domain shifts, and out-of-scope requests.

Out-of-Scope Requests

This release is a closed-set 10-class intent classifier.

It always assigns an input to one of the supported intents and does not provide a production-ready out-of-scope (OOS) detector.

Experimental confidence-margin and embedding-distance approaches were evaluated during development, but they are not part of the v0.1 inference pipeline.

Applications using this model should therefore not interpret a predicted intent as proof that the input belongs to the supported business scope.

Limitations

  • The model supports only the 10 intents listed above.
  • Training and evaluation data are not representative of all real-world Turkish contact-center conversations.
  • The dataset includes synthetic examples.
  • Real ASR transcripts may contain noise not sufficiently represented in the current evaluation data.
  • Semantically similar but unsupported business requests may be assigned to one of the supported intents with high confidence.
  • Confidence scores should not be interpreted directly as calibrated probabilities.
  • Additional evaluation is required before production deployment.

Example

Input:

Yeni hattımı nasıl aktif edebilirim?

Prediction:

hat_aktivasyonu

Input:

Modemim açılmıyor ve ışıkları yanmıyor.

Prediction:

cihaz_sorunu

Version

v0.1

Initial Hugging Face release of the V2.2 Turkish contact-center intent classifier.

Downloads last month
53
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yasinunal/turkish-contact-center-intent-classifier

Finetuned
(208)
this model