Multilingual Pizza Review Sentiment

Three-way sentiment for English and Malay pizza reviews, where the hard part is the middle.

Context
Deep Learning coursework, individual
My part
Everything: labelling, preprocessing, model comparison
Stack
Python, TensorFlow, Keras, LSTM, GRU
Year
2026
Pizza review sentiment analysis

0.902

Macro F1 on the test set; 0.866 on Mixed/Neutral, the hardest class

What it is
Reviews labelled Negative, Mixed/Neutral or Positive from their scores, in two languages, after quality checks and duplicate handling.
What I did
I prefixed each review with a language token, fit the tokenizer on training text only, and compared SimpleRNN, LSTM, GRU and stacked variants. L2, LayerNorm, class weights and augmentation were tried and dropped when they did not help.
Result
0.904 test accuracy. Most remaining errors are short, vague or genuinely mixed reviews: a limit of small, score-derived labels, not something more tuning fixes.