All projects
No. 9
FakeNewsDetector AI
News classifier with DistilBERT and heuristics
- My role
- Sole author
- Period
- Mar 2026
The problem
The hardest misinformation to catch isn't obvious clickbait but text that mimics the language of science: "researchers from a European university", "the study hasn't been published yet". A classifier trained only on obvious fake news tends to label that kind of text as reliable.
What I built
I built it alone and published it on March 30, 2026.
- Data preparation: 5,000 articles per class from the ISOT dataset on Kaggle (
Fake.csvandTrue.csv), with title and body joined and cut to 1,800 characters. - A DistilBERT fine-tune with the Hugging Face Trainer as a 2-class classifier, REAL and FAKE, with early stopping.
- Heuristic rules: pseudoscience patterns (total protection, unpublished study, unnamed university), alarm signals and signals of a verifiable source.
- Fusion of the model and the rules into 3 labels: Reliable, Doubtful and Fake. Doubtful shows up when the model can't separate the 2 classes well, and 2 or more pseudoscience patterns force Fake.
- A Gradio interface that explains the verdict. Without a local model it uses public Hugging Face models and, as a last resort, the rules alone.
Architecture
Results
I don't publish an accuracy figure. The README reports a very high accuracy, but I don't consider it valid for two reasons I explain below: it's measured on the same set used to pick the model, and the dataset gives away the answer. An honest evaluation needs a separate, clean test set, and that's still pending.
Known limits
- The README says the model is RoBERTa, but
train.pyfine-tunes DistilBERT. The README is wrong. - The evaluation set is the same one that drives early stopping and picks the best checkpoint, so the metric is optimistic.
- In ISOT the real articles carry a "(Reuters)" dateline and the code doesn't strip it: the model can learn to recognize the news agency instead of truthfulness.
- The trained model isn't in the repository, and the README mentions a notebook and a CSV of pseudoscience examples that aren't there either.
- One of the fallback models is a sentiment analysis model (SST-2), not a news model.
- The dataset is English only and there are no tests.