A Statistical Profile of Linguistic Differences Between Objective and Subjective Sports Articles

Authors

  • Luis Eduardo Muñoz Guerrero Universidad Tecnológica de Pereira

DOI:

https://doi.org/10.67145/ks.v7i1.4134

Keywords:

subjectivity classification, sports journalism, text mining, linguistic features, descriptive statistics, natural language processing

Abstract

Automated subjectivity detection in sports journalism has been studied since the early 2010s, yet the dataset most closely associated with this task — the Sports Articles for Objectivity Analysis dataset, released by Rizk and Awad in 2018 through the UCI Machine Learning Repository — remains largely unexplored beyond the work of its original creators. This paper presents a descriptive statistical analysis of the 59 syntactic and semantic features included in the dataset (N = 1,000 articles), comparing their distributions between articles labeled objective (n = 635) and subjective (n = 365). Non-parametric hypothesis testing (Mann-Whitney U) together with effect size estimation (rank-biserial correlation) identifies which linguistic markers differ significantly between the two classes, using Benjamini-Hochberg correction for multiple comparisons.

Of the 59 features tested, 53 (89.8%) show statistically significant differences. However, we identify an important confound: subjective articles are, on average, nearly twice as long as objective ones (1,006.6 vs. 519.2 words), which inflates the apparent effect size of virtually every raw frequency-count feature. After re-testing all features as length-normalized rates, 48 features (81.4%) remain significant, with possessive pronoun density, second-person pronoun usage, and imperative verb density emerging as the most robust discriminators once length is controlled for.

A logistic regression classifier trained on the statistically significant features, evaluated with nested 5-fold stratified cross-validation (feature selection re-run within each training fold to avoid data leakage), achieves 82.3% accuracy (F1 = 0.732) using raw features and 83.5% accuracy (F1 = 0.767) using length-normalized features, both well above the 63.5% majority-class baseline (a paired test across folds found this raw-vs-normalized gap statistically indistinguishable from noise, p = .19). A length-only control model (using article word count as the sole predictor) reaches 72.3% accuracy on its own, indicating that roughly 45% of the full model's improvement over baseline is attributable to article length alone rather than genuine per-word stylistic signal.

We further find severe multicollinearity among the raw features (66% with VIF ≥ 5, including several apparently redundant feature pairs), which is substantially reduced by length normalization and which cautions against interpreting individual logistic-regression coefficients as independent effects. These findings offer an accessible, non-technical reference point for researchers considering this dataset, and demonstrate that a resource with almost no citation history can still support a small, self-contained, and reproducible empirical contribution — including a genuine methodological caveat (the length confound) that prior classification-focused studies on this dataset did not report.

Downloads

Published

2019-03-22

How to Cite

Luis Eduardo Muñoz Guerrero. (2019). A Statistical Profile of Linguistic Differences Between Objective and Subjective Sports Articles. Kurdish Studies, 7(1), 121–134. https://doi.org/10.67145/ks.v7i1.4134

Issue

Section

Articles