All explorations
Exploration

NLP · Machine learning · Domain adaptation

Active Learning

Amazon → Twitter domain adaptation and active data selection.

Status
Academic group project
Stack
Python · pandas · NumPy · NLTK · scikit-learn · TF-IDF · Logistic Regression · K-Means · Entropy Sampling · Margin Sampling · Density · Diversity · Max Distance · hybrid strategy · Matplotlib · Seaborn

01

Problem studied

A sentiment classifier trained on Amazon reviews is transferred to tweets — a shorter, more informal domain with a different vocabulary. The goal is to maintain a reasonable performance level on this new domain while limiting the annotation budget required.

  • Model trained on Amazon reviews
  • Transfer to tweets — shorter, informal data with a different vocabulary
  • Goal: maintain performance within a limited target annotation budget

02

Baseline and domain shift

The baseline combines standard NLP preprocessing, unigram/bigram TF-IDF vectorization, and logistic regression across three classes — positive, negative, neutral. Applied without adaptation to the Twitter domain, it loses most of its performance.

  • Amazon → Amazon: about 79% accuracy
  • Amazon → Twitter (zero-shot, no adaptation): about 45% accuracy
Confusion matrix of the baseline evaluated on the Amazon source domain.
Baseline on the Amazon source domain (~79% accuracy).
Confusion matrix of the baseline applied without adaptation to the Twitter domain.
Same model applied without adaptation to Twitter (~45% accuracy).

03

Active learning

Rather than annotating the target domain at random, an active learning strategy selects, at each budget step, the tweets deemed most useful to annotate in order to adapt the model.

  • Uncertainty — Entropy Sampling, Margin Sampling
  • Diversity / representativeness — Diversity (K-Means), Density, Max Distance
  • Hybrid — normalized combination of Entropy and Margin

04

Strategy comparison

Uncertainty-based strategies are the most effective in the early annotation budgets. Margin and Entropy clearly outperform random selection, while Density is much less suited to this domain. The different strategies gradually converge as the budget increases.

Accuracy on the Twitter domain by active-selection strategy and annotation budget.
Comparison of active-selection strategies across annotation budget.

05

Focus on Margin

Margin Sampling stands out as the most interesting strategy for the early budgets: performance climbs quickly with few annotated examples, then gradually plateaus. This confirms the value of selecting informative examples rather than annotating at random.

Learning curve of the Margin Sampling strategy as a function of annotation budget.
Margin Sampling learning curve: quick gains followed by gradual saturation.

06

Qualitative analysis

The mathematical choice of a strategy concretely changes the nature of the annotated examples. Margin and Entropy favor more ambiguous, information-rich tweets, useful for shifting the decision boundary, while Density tends to select shorter, more redundant messages — which explains its weak performance on this noisy domain.

07

Limits

The TF-IDF representation remains heavily dependent on the vocabulary observed in the source domain. Twitter’s vocabulary differs sharply from Amazon’s, and a significant share of the target domain’s terms stay out-of-vocabulary, which mechanically caps performance. The conclusions remain limited to the datasets and experimental protocol used.

Back

All explorations