Journal of the College of Physicians and Surgeons Pakistan
ISSN: 1022-386X (PRINT)
ISSN: 1681-7168 (ONLINE)
Affiliations
doi: 10.29271/jcpsp.2026.08.969ABSTRACT
Objective: To determine gender using cranial measurements from CT scans and to evaluate their predictive accuracy for gender classification.
Study Design: A cross-sectional study.
Place and Duration of the Study: Department of Forensic Medicine and Toxicology, Ziauddin University, Karachi, Pakistan, from July to December 2025.
Methodology: A total of 280 CT scans of adult male and female were used for fourteen cranial measurements, including foramen magnum length (sagittal diameter; FML) and foramen magnum breadth (transverse diameter; FMB), right and left frontal sinus height (RTFSH, LTFSH), right and left frontal sinus breadth (RTFSB, LTFSB), right and left frontal sinus depth (RTFSD, LTFSD), frontal bone inclination angle (FI), maximum cranial length (MCL) and maximum cranial breadth (MCB), interorbital breadth (IOB), lateral wall interorbital breadth (LWIOB), and bizygomatic breadth (BZB). Right and left frontal sinus index (RTFSI, LTFSI) and cephalic index (CI) were calculated. SPSS version 27 was used to calculate mean ± SD, compare means (independent samples t-test or Mann-Whitney U test), and estimate gender prediction accuracies using binary logistic regression (BLR), decision tree (DT), and ROC curve analyses.
Results: MCL, MCB, RTFSB, LTFSB, RTFSD, LTFSD, FI, FML, FMB, LWIOB, BZB, and LTFSI showed significant gender differences. MCL, MCB, FML, LWIOB, BZB, and FI were the significant gender predictors in the BLR model. ROC curve analysis with male gender as the state variable showed an AUC >0.5 for 13 measurements, except for FI. BZB was the most important classifier in DT analysis, followed by MCL. Assuming male gender as the positive class, the BLR model showed higher specificity (0.814 > 0.714) but lower sensitivity (0.764 <0.835) than the DT model.
Conclusion: All the cranial measurements showed significant gender differentiation except IOB, RTFSH, and LTFSH. BLR and DT models both can be employed for gender prediction.
Key Words: Craniometry, Gender determination, CT images, Binary logistic regression, Receiver operating characteristic curve analysis, Decision tree.
INTRODUCTION
Gender determination constitutes a fundamental step in biological profiling in forensic identification, disaster victim identification, and anthropological investigations.1-3 Accurate determination of gender significantly narrows the pool of possible identities. While the pelvis is regarded as the most sexually dimorphic skeletal structure,4 it is frequently un- available or damaged in forensic contexts, necessitating reliance on the skull as an alternative and robust indicator of gender.3,5
The human skull exhibits pronounced sexual dimorphism resulting from genetic and hormonal influences on craniofacial growth.5 Skull size,6 foramen magnum,7 frontal sinus,1,8 bizygomatic distance9 have also shown sexual dimorphism. However, the sensitivity and specificity of the cranial measurements in gender determination vary across populations, underscoring the need for population-specific standards. Computed tomography (CT)-based cranial morphometric studies have been conducted in the Pakistani population;6 however, gender determination using combined statistical and machine-learning approaches remains limited.
CT has emerged as a powerful tool for precise, reproducible, and non-invasive acquisition of cranial measurements.1,2 Metric approaches to gender determination have evolved from uni- variate comparisons to multivariate techniques, including discriminant function analysis (DFA) and regression models to enhance classification accuracy.5,9,10 Receiver operating characteristic (ROC) curve analysis is employed to evaluate the diagnostic performance of these techniques in terms of sensitivity, specificity, and optimal cut-off values for gender determination.11
Recently, machine learning techniques, including decision tree (DT) models, have gained attention in forensic science to complement traditional statistical approaches.12
Nevertheless, there remains a need for a well-structured study design that can systematically integrate classical statistical testing, multivariate modelling, diagnostic accuracy assessment, and machine learning techniques within a unified analytical framework to investigate the reliability and accuracy of cranial measurements for gender deter- mination. This study aimed to evaluate whether cranial measurements from CT scans of adult skulls can be reliably used for gender determination. It was hypothesised that cranial measurements show significant gender-related differences.
METHODOLOGY
This cross-sectional study was conducted at the Department of Forensic Medicine and Toxicology, Ziauddin University, Karachi. CT scans were collected from the Department of Radiology, Ziauddin University, Clifton, North Nazimabad and Kemari Campuses, from July to December 2025. After obtaining informed consent, participants were enrolled using purposive convenience sampling. Normal healthy adult males and females within the age range of 20 to 80 years with no pathology/injury/disease of the skull were included in the study. Adults with systemic diseases (e.g., hyper or hypothyroidism), syndromes (e.g., Down’s synd-rome), congenital developmental abnormalities, nutritional diseases affecting bones or who have undergone orthodontic treatment were excluded from the study.
CT images were obtained using the bone window settings of the scanning system (Toshiba Aquilion 16, 200 VAC, 60 kW generator, 7.5 MHU tube) with a slice thickness of 2 mm, a matrix of 512 × 512 pixels. Thirteen linear (in mm) and one angular measurement (o) were recorded (Figure 1). Landmarks were defined, and measurements were taken by two independent observers (the primary investigator and a radiologist) to ensure reliability.
Fourteen cranial measurements were obtained: maximum cranial length (MCL), maximum cranial breadth (MCB), right frontal sinus height (maximum) (RTFSH), right frontal sinus breadth (maximum) (RTFSB), right frontal sinus depth (maximum) (RTFSD), left frontal sinus height (maximum) (LTFSH), left frontal sinus breadth (maximum) (LTFSB), left frontal sinus depth (maximum) (LTFSD), frontal bone inclination angle (FI), interorbital breadth (IOB), lateral wall interorbital breadth (LWIOB), bizygomatic breadth (BZB), foramen magnum length (sagittal diameter; FML), and foramen magnum breadth (transverse diameter; FMB). Three indices (cephalic index (CI), right frontal sinus index (RTFSI), and left frontal sinus index (LTFSI) were also calculated: CI = (MCB ÷ MCL) × 100 and FSI = (FSH ÷ FSB) × 100.
Sample size was calculated by using the means (μ) of the cranial measurements using the following formula:

where n = sample size of each gender, ʊ= population variance, μ1 and μ2 = mean measurement of maximum length of skull in both genders, α = significance level 0.05 (z value = 1.96), and 1- β = power of the test (z value = 0.842).
The mean (μ) values of the maximum length of the skull of both male and female were taken as μ1 (173.86 ±6.65) and μ2 (165.78 ± 7.19), respectively, from a research study.3 Minimum sample size was calculated as 67 samples of each gender.
Binary logistic regression (BLR) was used to predict gender (outcome variable) for the set of 14 cranial measurements (predictor variables), which needs to have at least 10 EPV (events per variable) to ensure model stability by reducing bias and preventing overfitting, i.e. recognising noise in the dataset as a generalisable pattern.14 Therefore, 140 individuals from each gender were required. The present study included 280 adult subjects (140 males and 140 females), thereby exceeding the minimum required sample size to ensure adequate power for both univariate (independent samples t-test) and multivariate analyses (BLR and ROC curve analysis), while minimising the likelihood of Type I and Type II errors.
Data were analysed using SPSS version 27 (IBM Corp., Armonk, NY, USA). Descriptive statistics (mean ± SD) were calculated. The Shapiro-Wilk test of normality was used for each cranial measurement. The independent sample t-test or the Mann-Whitney U test was used to compare the means of each cranial measurement. Cohen’s d or Pearson’s r were measured to see the effect size.
The univariate BLR model estimates the odds of male gender associated with each measurement. Multivariate BLR was used to predict gender (outcome variable) using all 14 cranial measurements as predictor variables. An equation was derived for each variable, and application of the equation to the corresponding variable value yielded a predicted value. The cut point to predict gender was set to 0.5. The equation provided by the model is: p(Male) = 1 / (1 + e−Logit), where e is the mathematical constant 2.718, and Logit = B₀ + B₁X₁ + B₂X₂ (B₀ = intercept/constant, B = regression coefficient for each cranial measurement, and X = observed cranial measurement). Predicted probabilities >0.5 suggest that the crania are more likely male, while probabilities <0.5 indicate the crania are more likely female. ROC curve analysis was used to estimate the diagnostic/predictive accuracy of each cranial measurement by quantifying the area under the curve (AUC) values with corresponding sensitivity and specificity.11
Classification and Regression Tree (CRT)-based DT analysis was conducted to develop practical classification rules and derive cranial measurement threshold values for gender determination.12 The split sample method was used to validate the model with 60% training sample and 40% test sample of the dataset. CRT splits the dataset into binary nodes based on the optimal cut-off of predictor variables. The CRT algorithm selects the optimal split using the Gini index criterion. Minimum parent and child node sizes were empirically optimised at 30 and 15, respectively. This approach maintains an adequate number of observations within each node, thereby ensuring stable classification performance relative to the study sample size. All statistical tests were two- tailed, and p <0.05 was considered statistically significant.
Both BLR and DT models were evaluated for their performance as diagnostic tests for male gender using sensitivity and specificity values calculated from classification tables, as well as the ROC analysis curve.
RESULTS
Two hundred and eighty CT scans of the skull were measured with a male-to-female ratio of 1:1. The mean age of the participants was 44.48 ± 17.45 years for males and 48.85 ± 17.93 years for females.
Descriptive statistics showed that all 17 measurements were greater in males than in females (Table I). Of which 12 measurements were significantly different between the two genders (<0.05); BZB showed the largest effect size (r = 0.524).
Univariate BLR was performed to identify cranial measurements that independently predicted gender classification. Cranial measurements with p <0.05 were considered significant predictors and highlighted in bold (Table II).
Table I: Comparison of cranial measurements between genders.
|
Cranial measurements (only abbreviations were given) |
Gender |
Mean/median |
SD/IQR |
Cohen’s d/ Pearson’s r†† |
p-values |
|
a. MCL |
Female |
165.49 |
7.248 |
-0.718 |
<0.001* |
|
Male |
171.21 |
8.596 |
|||
|
b. MCB† |
Female |
132.83 |
5.697 |
0.344 |
<0.001** |
|
Male† |
137.30 |
7.75 |
|||
|
c. RTFSH |
Female |
26.00 |
8.409 |
-0.142 |
0.235* |
|
Male |
27.24 |
8.986 |
|||
|
d. RTFSB |
Female |
22.94 |
8.323 |
-0.281 |
0.020* |
|
Male |
25.52 |
9.942 |
|||
|
e. RTFSD |
Female |
9.11 |
2.555 |
-0.525 |
<0.001* |
|
Male |
10.63 |
3.190 |
|||
|
f. LTFSH |
Female |
26.12 |
7.570 |
-0.232 |
0.054* |
|
Male |
27.98 |
8.463 |
|||
|
g. LTFSB† |
Female |
23.78 |
9.160 |
0.189 |
<0.001** |
|
Male† |
28.75 |
17.00 |
|||
|
h. LTFSD† |
Female |
9.37 |
2.926 |
0.280 |
<0.001** |
|
Male† |
11.05 |
4.07 |
|||
|
i. FI |
Female |
77.19 |
7.179 |
0.549 |
<0.001* |
|
Male |
73.25 |
7.176 |
|||
|
j. IOB† |
Female† |
19.50 |
2.58 |
0.086 |
0.149** |
|
Male† |
19.94 |
3.55 |
|||
|
k. LWIOB† |
Female† |
94.55 |
4.40 |
0.419 |
<0.001** |
|
Male† |
98.40 |
5.65 |
|||
|
l. BZB† |
Female† |
122.5 |
4.00 |
0.524 |
<0.001** |
|
Male |
128.68 |
5.721 |
|||
|
m. FML† |
Female |
36.25 |
4.30 |
0.372 |
<0.001** |
|
Male |
38.61 |
3.482 |
|||
|
n. FMB |
Female |
31.13 |
3.266 |
-0.527 |
<0.001* |
|
Male |
32.80 |
3.042 |
|||
|
Female |
80.45 |
5.452 |
-0.029 |
0.810* |
|
Male |
80.28 |
5.885 |
|||
|
Female† |
117.15 |
50.96 |
-0.088 |
0.141** |
|
Male† |
110.18 |
39.50 |
|||
|
Female† |
114.28 |
43.28 |
-0.171 |
0.004** |
|
Male† |
100.98 |
41.4 |
|||
|
†Cranial measurements showed significant (p <0.05) on the Shapiro-Wilk test; therefore, they were presented as the median and interquartile range (IQR), and the effect size was calculated using Pearson’s r. ††Cohen’s d effect size was interpreted as small (0.20 to 0.49), medium (0.50 to 0.79), and large (≥0.80)15 and Pearson’s r effect size as small (≥0.1), medium (≥0.3), and large (≥0.5).15 *p-value of independent sample t-test. These cranial measurements showed a non-significant Shapiro-Wilk test (p >0.05). **p-value of the Mann-Whitney U test. These cranial measurements showed a significant Shapiro-Wilk test (p <0.05). |
|||||
Table II: Univariate BLR analysis of cranial measurements for gender prediction.
|
Cranial measurements (only abbreviations were given) |
B (regression coefficient) |
B0 (constant/ intercept) |
Wald statistics (p-values*) |
Prediction accuracy (male)% |
Prediction accuracy (female)% |
|
|
a |
MCL |
0.091 |
-15.394 |
28.656 (<0.001) |
64.3 |
62.1 |
|
b |
MCB |
0.108 |
-14.565 |
25.994 (<0.001) |
67.9 |
64.3 |
|
c |
RTFSH |
0.016 |
-0.437 |
1.405 (0.236) |
53.6 |
52.1 |
|
d |
RTFSB |
0.031 |
-0.747 |
5.348 (0.021) |
52.9 |
57.9 |
|
e |
RTFSD |
0.185 |
-1.823 |
16.906 (<0.001) |
59.3 |
63.6 |
|
f |
LTFSH |
0.029 |
-0.768 |
3.68 (0.055) |
55.0 |
60.0 |
|
g |
LTFSB |
0.041 |
-1.054 |
10.639 (0.001) |
58.6 |
61.4 |
|
h |
LTFSD |
0.190 |
-1.955 |
19.739 (0.001) |
60.7 |
62.1 |
|
i |
FI |
-0.076 |
5.745 |
18.531 (<0.001) |
65.0 |
63.6 |
|
j |
IOB |
0.082 |
-1.606 |
3.199 (0.074) |
45.0 |
61.4 |
|
k |
LWIOB |
0.154 |
-14.797 |
23.541 (<0.001) |
67.9 |
67.9 |
|
l |
BZB |
0.241 |
-30.223 |
54.294 (<0.001) |
70.0 |
81.4 |
|
m |
FML |
0.247 |
-9.195 |
34.787 (<0.001) |
65.7 |
66.4 |
|
n |
FMB |
0.170 |
-5.423 |
17.027 (<0.001) |
64.3 |
55.0 |
|
*p-values were calculated using the Wald test in BLR. |
||||||
Table III: Multivariate BLR analysis of cranial measurements for gender prediction.
|
1. Nagelkerke R² |
0.515 |
|||
|
2. Hosmer–Lemeshow goodness-of-fit test (p-value) |
0.103 |
|||
|
3. Prediction accuracy (predicted/ observed) in male participants (true positive† |
107/140 (76.4%) |
|||
|
4. Prediction accuracy (predicted/ observed) in female participants (true negative† |
114/140 (81.4%) |
|||
|
5. Sensitivity (true positive/ (true positive + false negative) |
107/107 + 33 = 0.764 |
|||
|
6. Specificity (true negative/ (true negative + false positive) |
114/114 + 26 = 0.814 |
|||
|
Cranial measurements (only abbreviations were given) |
B (regression coefficient) |
Wald statistics |
p-values* |
Odds ratio (OR) or Exp(B) |
|
a. MCL |
0.059 |
7.026 |
0.008 |
1.061 |
|
b. MCB |
0.080 |
6.905 |
0.009 |
1.084 |
|
c. RTFSH |
0.002 |
0.004 |
0.952 |
1.002 |
|
d. RTFSB |
-0.020 |
0.558 |
0.455 |
0.981 |
|
e. RTFSD |
-0.009 |
0.013 |
0.910 |
0.991 |
|
f. LTFSH |
-0.027 |
0.758 |
0.384 |
0.974 |
|
g. LTFSB |
0.005 |
0.045 |
0.832 |
1.005 |
|
h. LTFSD |
0.112 |
2.158 |
0.142 |
1.118 |
|
i. FI |
-0.057 |
5.642 |
0.018 |
0.945 |
|
j. IOB |
0.000 |
0.000 |
0.994 |
1.000 |
|
k. LWIOB |
0.069 |
4.545 |
0.033 |
1.072 |
|
l. BZB |
0.118 |
8.012 |
0.005 |
1.125 |
|
m. FML |
0.178 |
9.642 |
0.002 |
1.195 |
|
n. FMB |
0.017 |
0.078 |
0.779 |
1.017 |
|
Constant (intercept) |
-45.217 |
36.863 |
0.000 |
0.000 |
|
†Male gender was taken as the positive class and female as the negative class in calculating the specificity and sensitivity of the BLR model. *p-values were calculated using the Wald test in BLR. |
||||
Following univariate BLR, multivariate BLR was performed to identify cranial measurements that were significant predictors of gender after adjusting for the potential influence of other cranial measurements (Table III).
BLR showed a good fit to the present study dataset, as the Hosmer–Lemeshow goodness-of-fit test was non-significant (p >0.05, Table III). The Nagelkerke pseudo-R² value indicated that 51.5% of the variation in cranial measurement data was explained by the BLR model. Prediction accuracy for females was higher than for males. The prediction performance of six cranial measurements (MCL, MCB, FML, LWIOB, BZB, and FI) was significant (p <0.05).
DT analysis produced a three-layered, five-node model for training (n = 167) and test (n = 113) samples, with BZB as the primary splitter and MCL as the secondary splitter (Figure 2). A BZB value >125.55 mm classified samples as male in both the training (76.1%) and test (82.5%) sets. In contrast, BZB values ≤125.55 mm combined with MCL values ≤172.25 mm classified samples as female in both the training (84.6%) and test (75.6%) sets. Assuming male sex as the positive class, the DT model demonstrated a sensitivity of 0.835 and a specificity of 0.714 (Figure 2).
ROC Curve analysis (Figure 3A) showed AUC values for 13 cranial measurements ranging from 0.550 (RTFSH, IOB) to 0.803 (BZB), with ROC curves above the reference line, except for FI (AUC = 0.344). Nonetheless, FI was identified as a significant predictor in BLR with β = –0.057, p = 0.018 (Table III), indicating an inverse relationship with the outcome (male gender). ROC curve analysis and BLR findings were suggestive of higher FI values positively associated with female gender.
Comparative analysis of the two models, BLR and DT, was performed using two methods. The first method was to compare their sensitivity and specificity. The specificity and sensitivity of the DT model were calculated using the values of true and false gender identification cases from both the training and test sets, as illustrated in Figure 2.
Figure 1: Measurements shown on CT scans of human skull.
When compared with the BLR model (Table III), the speci-ficity of BLR was higher than that of the DT model, while BLR sensitivity was lower than that of the DT model. The second method was used to compare their AUC in the ROC curve analysis. Predicted probabilities were calculated from the BLR and DT models and used as test variables for ROC analysis. The AUC values for the BLR and DT models were 0.867 and 0.797, respectively, indicating that both models had strong predictive performance for gender determination (Figure 3B). However, the ROC curve of BLR was smoother than the stepwise curve of DT, indicating that the DT model has fewer probability thresholds than the BLR model. In ROC curve analysis, sensitivity and specificity varied across diffe-rent probability thresholds and were therefore not identical to those obtained by the first method, which was based on a fixed threshold (i.e., 0.5).
DISCUSSION
Sexual dimorphism in humans is the product of environmental variations and genetic factors. Gender and ethnicity are the primary determinants of skeletal measurements.15 However, the performance accuracies of cranial measurements for gender differentiation vary in terms of prediction, classification, sensitivity, and specificity.
Prediction accuracies of FML were reported 74.4% (male) and 64.4% (female), while FMB accuracies were 69.8% (male) and 66.7% (female) using the discriminant analysis.16 BLR and ROC curve analyses showed prediction accuracies of 69.6% for FML and 66.4% for FMB,7 which supported the finding of the present study. Hence, the foramen magnum appeared as a significant predictor of gender. The difference in prediction accuracies may be attributed to the methodology used, such as the use of a digital calliper for measurements,7 the use of different statistical models,16 or variations in sample size.7,16
Accuracy of gender determination for LWIOB (also described as bi-orbital or extra-orbital distance in some studies) using regression and ROC analyses ranged from 74%17 to 84%18, which strongly corroborated the findings of the present study. Difference in sample sizes (n = 100 vs. n = 200) may account for the wide variation in prediction accuracies for LWIOB.17,18
Figure 2: DT models for the training and test sets.
Figure 3: (A) ROC and AUC values (shown in parentheses) for the BLR model for each cranial measurement. (B) ROC curves and AUC values (shown in parentheses) for the BLR and DT models.
Similarly, 79.7% accuracy was reported for the FS.19,20 In the present study, FSB and FSD were found to be more discriminatory than FSH, and the left sinus was more discriminatory than the right. This finding was also supported by research studies19-20. Although sinus height was also reported to be discriminatory.20,21
For FI, accuracies were reported to range from 66% in the Chinese population to 81% in Portuguese, White American, and Black American populations,22 while in the present study, FI accuracies were 65.0% for males and 63.6% for females, indicating greater similarity to the Chinese population. Petaros et al. reported a 5.8° difference in FI between females and males (79.1° vs. 73.3°) using 2D and 3D photographs of 122 dry skulls from the Croatian population, compared with a 3.94° difference in the present study (77.19° vs. 73.25° for females and males, respectively).23 Difference in methodology may account for variation in the results. However, both studies reported more vertical FI in females than in males.22,23
BZB was reported to be sexually dimorphic, showing an effect size of 1.39 in the US white population (n = 6068)9 compared with 1.167 in the present study. Kalan et al.'s results strongly corroborated the findings of the present study, reporting BZB and MCL as the most accurate gender predictors.24 Although MCL and MCB showed significant gender differentiation in the present study, CI did not show a significant difference, and both genders showed the mesocephalic skull, as did their neighbouring populations.25
The variation in performance accuracies emphasises the need for an exhaustive search for strong determinants of sex using more accurate cranial measurements and comprehensive statistical analyses to identify models that can more explicitly explain gender-based differences in cranial measurements. Fewer studies have employed machine learning approaches for craniometric studies.12 Traditional approaches, such as BLR and DFA, were commonly used for such studies.5,7,14,16,18 Employing data science and machine learning algorithms may strengthen the study findings and validate their forensic application.
This study employed a comprehensive and novel four-tier analytical framework consisting of univariate and multivariate BLR, ROC curve analysis, and DT analysis, applied to 14 cranial measurements within a single cohort, together with an ROC-derived comparison of the AUC values of the BLR and DT models.
This study is limited by a lack of inter-observer reliability testing and a moderate sample size from a single institution, which may restrict generalisability to broader populations. Larger, multi-centre, and ethnically diverse datasets using advanced machine learning algorithms with external validation are recommended to enhance the robustness of gender estimation models.
CONCLUSION
Cranial measurements are good predictors of gender differentiation; however, BZB, FML, LWIOB, MCL, MCB, and FI outperformed the others and can be preferred in establishing a valid statistical model for gender determination.
ETHICAL APPROVAL:
Ethical approval was obtained from the Ethical Review Committee of Ziauddin University, Karachi, Pakistan (Ref Code: 10030425NAFOR; dated: 9 April 2025).
PATIENTS’ CONSENT:
Informed consent was obtained from each patient, and anonymity was assured.
COMPETING INTEREST:
The authors declared no conflict of interest.
AUTHORS’ CONTRIBUTIONS:
NAA: Conceptualisation, literature review, data collection, statistical analysis, write-up, critical review, and editing.
QH: Conceptualisation, critical review, critical feedback, and proofreading.
Both authors approved the final version of the manuscript to be published.
REFERENCES