Does Model Choice Matter in Low-Resource, Code-Mixed Crisis Discourse? A Controlled Comparison of Five Sentiment-Analysis Architectures
Abstract
Most sentiment-analysis models are developed and validated on high-resource domains where sentiment is fairly unambiguous (product reviews, generic tweets). It is far less clear how these architectures behave on low-resource, code-mixed, emotionally charged language, such as social-media discourse during a prolonged crisis. This gap matters in practice: governments, NGOs, and journalists increasingly rely on automated sentiment tools to monitor unfolding crises in exactly these harder conditions, often without knowing whether a given model is reliable enough for that purpose. We use an eight-year socio-political crisis in a Central African country as a case study to ask, concretely, whether model choice matters under these conditions, and why. We collected posts and comments from Twitter/X and Facebook discussing the crisis and cleaned the text (tokenization, stemming, stop-word removal). Off-the-shelf VADER labelling performed poorly on crisis-specific language for instance, treating calls for a detained leader's release as neutral rather than as an expression of grievance so we built a customized variant that layers domain-specific keyword and phrase rules onto VADER's lexicon before using the resulting labels to train five model families spanning three generations of NLP: classical learners (Naive Bayes, SVM) over TF-IDF/bag-of-words features; a CNN and an RNN over learned embeddings; and a fine-tuned BERT model. Each model was evaluated with k-fold cross-validation. Because labels generated by a rule-based method can be easier for a model to reproduce than genuine sentiment, we treat raw accuracy with caution and prioritize precision, recall, F1, and AUC-ROC, given the class imbalance in the data. Across those metrics, the sequence-aware models (RNN, BERT) substantially outperform the classical baselines and the CNN in this setting. The CNN scores lowest of the five; we interpret this as consistent with its limited capacity to model the long-range, code-mixed dependencies that dominate crisis narratives, though we did not run a targeted ablation to confirm that mechanism, and note it here as an interpretation rather than a demonstrated cause. The RNN edges out BERT on aggregate F1 and reaches 99% raw accuracy, a figure we report as an upper bound shaped by the weak-supervision labels rather than as a clean ground-truth result. BERT's attention weights, while not a formal explainability method, do provide interpretable token-level evidence for individual predictions that the other models do not offer. A time-resolved reading of the sentiment stream also shows a sharp, dateable spike in negative sentiment coinciding with a specific violent incident during the crisis, external evidence that the pipeline is tracking a real signal rather than noise. The main contribution is not simply that five models were compared, but what the comparison suggests: sequence-aware architectures appear particularly important for handling the ambiguity and code-mixing typical of crisis discourse, since they substantially outperform the non-sequence baselines tested here a pattern that should be checked on other crisis-discourse datasets before it is generalized further. We also surface a methodological risk specific to this setting a rule-based labeler tuned to the same vocabulary the model later learns from can inflate apparent performance and flag it explicitly rather than let the headline accuracy stand unqualified; the current results have not yet been validated against an independently human-annotated holdout, and doing so, together with a temporal (rather than random) train/test split, is the immediate next step. Together, these findings extend the sentiment-analysis and opinion-mining literature (Wankhade et al., 2022) and the underlying model architectures (Hochreiter & Schmidhuber, 1997; Devlin et al., 2019) to a domain where robustness, not headline accuracy, determines whether such a tool is actually usable by journalists, NGOs, or policymakers.