Part 5: Artificial Intelligence for Detecting SIM Boxing: Machine Learning, Graph Analytics, and Anomaly Detection
Automated detection of SIM boxing
Abstract
Parts I through IV of this series established the technical mechanics of SIM boxing, its treatment as evidence in Kenyan litigation, and the layered regulatory-to-judicial process through which SIM boxing findings are made and reviewed. A recurring theme across Parts II and IV is that Kenyan regulators and tribunals have so far relied on comparatively simple inferential reasoning, most notably, the absence of expected international numbers in Call Detail Records to sustain SIM boxing findings, without the benefit of purpose-built detection systems.
This paper surveys the state of the art in automated SIM boxing detection, covering supervised machine learning classifiers trained on CDR-derived features, graph-based and network-analytic approaches that model subscriber relationships, and emerging physical-layer techniques that detect SIM boxes from cellular signalling anomalies rather than call records alone. It also considers the adversarial dimension of this problem: fraudsters actively adapt their infrastructure to evade detection, producing an ongoing technical arms race. The paper closes by considering what these detection methods could mean for Kenyan regulatory and forensic practice, including their implications for the standard of appellate deference discussed in Part IV, and sets up the digital forensic framework to be developed in Part VI.
1. Introduction
Part I of this series identified the volume, velocity, and rotating identifiers of SIM boxing traffic as the central obstacles to manual investigation, and suggested that these characteristics make SIM boxing a natural application for artificial intelligence and anomaly detection. Part IV then showed, through the Elige and Geonet decisions, that Kenyan regulatory and tribunal practice has so far relied on a comparatively narrow evidentiary technique: examining whether an operator’s Call Detail Records contain international numbers consistent with the operator’s own account of how its traffic originates, and drawing an adverse inference from their absence. That technique proved legally sufficient to sustain a SIM boxing finding through two layers of appeal, but it is a far cry from the automated, feature-rich detection systems that have been developed in the wider telecommunications security literature over the past decade.
This paper surveys that literature, organised around three broad families of technique: supervised machine learning classifiers trained on features engineered from CDRs; graph-based and network-analytic methods that model the relationships between subscribers and numbers rather than treating each call record in isolation; and physical or signalling-layer techniques that detect SIM box hardware directly from the way it interacts with the cellular network, independent of call content or CDR analysis altogether. It also addresses the adversarial character of the problem: SIM box operators have adapted their infrastructure specifically to defeat CDR-based detection, which has driven researchers toward increasingly sophisticated countermeasures.
2. The Limits of Traditional Detection Methods
Historically, mobile network operators relied on two principal tools to detect SIM boxing: Test Call Generation (TCG), in which an operator originates test calls designed to travel through likely bypass routes and checks whether they are re-terminated locally, and rule-based Fraud Management Systems (FMS) that flag traffic matching pre-defined suspicious patterns, such as unusually high call volumes on a single SIM or unusually short average call durations. A substantial body of research has documented that both approaches are readily circumvented once fraudsters are aware of them: TCG numbers can be identified and blacklisted by SIM box operators, and static FMS rules can be evaded simply by adjusting traffic patterns to fall just outside the flagged thresholds.
The comprehensive 2021 survey of SIM boxing by Kouam, Viana and Tchana catalogues this pattern across the detection literature: as soon as a static detection signature becomes known, fraud operators adjust their behaviour to defeat it, which has pushed the field toward machine learning approaches capable of learning more complex, harder-to-evade behavioural signatures directly from data, and toward network- and hardware-level signals that are more costly for fraudsters to disguise.
3. Feature Engineering from Call Detail Records
Machine learning approaches to SIM boxing detection typically begin by engineering a feature set from raw CDRs that captures behavioural patterns distinguishing SIM box traffic from genuine subscriber traffic. Commonly used features include: the ratio of outgoing to incoming calls for a given number, since SIM boxes are overwhelmingly outgoing; the diversity of B-party (called) numbers dialled from a single A-party number, since a SIM box terminating high volumes of international traffic will typically call a much larger and more geographically dispersed set of numbers than an ordinary subscriber; average and total call duration; calling patterns across the day and week, since SIM boxes often operate on schedules dictated by international traffic volumes rather than typical human behaviour; and cell-site or location-area consistency, since a SIM card that never changes physical location in a manner consistent with normal mobility, yet generates very high call volumes, is characteristic of a SIM box rather than a handset carried by a person.
4. Supervised Machine Learning Classifiers
A substantial strand of the literature applies standard supervised classification algorithms to labelled CDR datasets, training models to distinguish known fraudulent numbers from known legitimate ones. Fitsum’s 2020 study of near-real-time SIM box detection at Ethio Telecom is illustrative of this approach in an African operator context directly comparable to Kenya. The study applied a sliding-window aggregation technique to reduce detection latency to approximately one hour, and compared three supervised classifiers: Random Forest, Artificial Neural Networks, and Support Vector Machines, finding that Support Vector Machine models achieved the highest classification accuracy among the algorithms tested on the operator’s CDR dataset.
Other studies in this tradition have applied neural network architectures, including multilayer perceptrons and recurrent architectures such as Long Short-Term Memory networks, to sequential CDR data, and rough or fuzzy set-based classifiers designed to handle the inherent uncertainty and incompleteness of telecommunications data. A common thread across this literature is that supervised classifiers can achieve high measured accuracy on historical, labelled datasets, but their performance depends heavily on the availability of reliable fraud labels for training, which are themselves typically derived from prior investigations of the kind discussed in Part IV creating a degree of circularity that automated systems inherit from the underlying regulatory and investigative process that labels the training data in the first place.
5. Graph-Based and Network-Analytic Approaches
A more recent strand of research treats SIM boxing detection as a graph problem rather than a per-record classification problem, on the reasoning that fraudulent numbers rarely act in isolation and instead form identifiable clusters or networks of related activity shared switching infrastructure, coordinated calling schedules, or SIM cards that are rotated across a common pool of hardware. Graph-based methods construct a network in which subscribers or numbers are nodes and calls are edges, then apply graph machine learning techniques, such as node embedding and graph neural networks, to classify nodes as fraudulent or legitimate based on their position and behaviour within the overall network structure.
A significant technical obstacle in applying graph methods to real-world CDR data is that genuine call graphs are often sparse, meaning most subscriber pairs never call each other directly, which limits the connectivity that graph algorithms rely on. Hu and colleagues’ 2022 Bridge To Graph (BTG) framework addresses this directly by reconstructing a denser graph from subscriber behavioural similarity before applying graph-based fraud classification, and reports improved detection performance over non-graph baseline methods on a real-world telecommunications CDR dataset. Related work on social network analysis of telecommunications data has also demonstrated the use of graph centrality measures to identify influential nodes within a subscriber network, and to detect multi-SIM subscribers operating across several accounts, both of which are relevant behavioural markers in a SIM boxing context.
6. The Adversarial Dimension: Evasion and the Detection Arms Race
SIM boxing detection is not a static classification problem; it is an adversarial one, in which fraud operators have both the incentive and the technical means to adapt their infrastructure specifically to defeat known detection techniques. Remote SIM card association, in which a SIM box physically separates its GSM radio hardware from the SIM cards it uses, distributing SIM cards across multiple genuine handsets or SIM farms connected to the box over IP is a leading example, and one already noted as a challenge for investigators in Part I of this series.
Kouam, Viana and Tchana’s 2024 study, examining the extent to which fraudsters can disguise SIM boxing traffic against detection systems, demonstrates empirically that fraud operators are capable of tuning their calling patterns to closely mimic legitimate subscriber behaviour once they have some knowledge of how detection systems operate, substantially degrading the effectiveness of CDR-based classifiers trained on older, more distinguishable fraud patterns. This finding has direct implications for the evidentiary reasoning discussed in Part IV: if fraud operators can and do adapt to defeat CDR-based signatures, then a regulator’s reliance on a single, well-known signal, such as the absence of international numbers, risks becoming decreasingly reliable as SIM boxing operators become aware of it and adjust accordingly.
7. Physical and Signalling-Layer Detection
Partly in response to this adversarial dynamic, the most recent research has moved beyond CDR analysis altogether, toward detecting SIM box hardware from properties of the cellular network attachment process itself, which are far more costly for fraudsters to disguise because they arise from physical and protocol-level constraints rather than from patterns in call behaviour that can simply be re-tuned.
The 2025 SigN system, developed by Kouam and colleagues, is a leading example of this approach. Rather than analysing CDRs, SigN detects the latency signature created when a SIM box’s remote SIM association architecture routes authentication signalling over the internet between a decoupled radio unit and its SIM cards, rather than through the tightly coupled hardware path of an ordinary handset. Through controlled indoor and outdoor experiments, the study found that SIM box devices produced attachment latencies markedly higher than standard devices, particularly during the network authentication phase, and attributed this overhead to structural features of LTE authentication and internet-based signalling that are not easily masked through behavioural tuning of the kind discussed in Section 6.
8. Implications for Kenyan Regulatory and Judicial Practice
Read against the findings of Part IV, this survey of the detection literature suggests several practical implications for CAK and for Kenyan digital forensic practice more broadly:
- The evidentiary technique actually used in the Geonet dispute examining CDRs for the presence or absence of a single expected feature is a legitimate but comparatively narrow slice of a much broader available toolkit, and CAK’s future determinations could draw on a richer, systematically engineered feature set without departing from the balance-of-probabilities standard the Tribunal and High Court have already endorsed.
- Because CDR-based signatures are known to be subject to adversarial adaptation, any automated detection system CAK deploys should be periodically revalidated against current fraud patterns rather than treated as a fixed, one-time methodology, and its outputs should be documented transparently enough to survive the kind of appellate scrutiny discussed in Part IV, notwithstanding the deference specialist regulators are typically afforded.
- Physical and signalling-layer detection methods, while more technically demanding to deploy, offer a form of evidence that does not depend on an operator’s own CDR submissions and could reduce future disputes’ dependence on inferences drawn from data the respondent operator itself controls and produces.
- Graph-based analysis across multiple operators’ interconnection data could help address the numbering-resource and transit-traffic issues that CAK found inconclusive in the Geonet determination for want of corroborating evidence, discussed in Part IV, by surfacing cross-operator patterns invisible within a single operator’s own records.
9. Conclusion
The detection techniques surveyed in this paper supervised classification, graph-based network analysis, and physical-layer signalling analysis represent a substantially more developed evidentiary toolkit than the single-feature CDR reasoning that has, so far, proven sufficient to sustain SIM boxing findings through Kenya’s regulatory and appellate process. That sufficiency should not be mistaken for optimality: the adversarial literature reviewed in Section 6 indicates that reliance on a narrow, well-known evidentiary signature carries a built-in expiry date as fraud operators adapt around it. The final paper in this series draws together the legal, evidentiary, and technical threads established across Parts I through V into a proposed digital forensic framework for how Kenyan investigators, regulators, and courts should collect, preserve, analyse, and present SIM boxing evidence going forward.
๐ References
Legislation
Case Law (carried forward from Part IV)
Part VI (digital forensic framework) forthcoming · series DOI: 10.5281/zenodo.xxxx
Comments
Post a Comment