Style Scalpel research guide
Style Scalpel Benchmark Results
Review Style Scalpel benchmark results, validation scope, dataset conditions, limitations and responsible interpretation of AI-authorship detection performance.
IDMGSP scientific-paper benchmark
Style Scalpel was evaluated against the IDMGSP benchmark for distinguishing human-written and machine-generated scientific papers. Values below are accuracy percentages published in the project repository. They describe specific scientific-paper dataset conditions and should not be generalized to every genre, model, or real-world setting.
| Model | Train dataset | TEST | OOD-GPT3 | OOD-REAL | TECG | TEST-CC |
|---|---|---|---|---|---|---|
| Style Scalpel / stylometric model | TRAIN | 98.68 | 17.60 | 97.52 | 100.0 | 14.67 |
| Style Scalpel / stylometric model | TRAIN-CG | 98.23 | 15.20 | 97.50 | 95.80 | 9.55 |
| Style Scalpel / stylometric model | TRAIN + GPT-3 | 98.79 | 98.90 | 95.65 | 100.0 | 23.18 |
Interpretation and limitations
The model is competitive on the main TEST split and strong on OOD-REAL and TECG under these conditions. OOD-GPT3 performance is weak when GPT-3 examples are absent from training, and TEST-CC remains weak. The project documentation notes that TEST-CC appears to contain substantial encoding or extraction artifacts.
Benchmark performance on specific datasets does not establish universal AI-detection accuracy. Results depend on domain, genre, model family, text length, editing, and other conditions.
These results are benchmark evidence, not a guarantee and not proof that a particular document has a particular origin.
Benchmark source
Abdalla, M. H. I., Malberg, S., Dementieva, D., Mosca, E., & Groh, G. (2023). A benchmark dataset to distinguish human-written and machine-generated scientific papers. Information, 14(10), 522. Read the benchmark paper.
View the complete comparison table in the project repository or read about the research behind Style Scalpel.