SSStyle ScalpelForensic authorship intelligence

Style Scalpel research guide

Style Scalpel Benchmark Results

Review Style Scalpel benchmark results, validation scope, dataset conditions, limitations and responsible interpretation of AI-authorship detection performance.

IDMGSP scientific-paper benchmark

Style Scalpel was evaluated against the IDMGSP benchmark for distinguishing human-written and machine-generated scientific papers. Values below are accuracy percentages published in the project repository. They describe specific scientific-paper dataset conditions and should not be generalized to every genre, model, or real-world setting.

ModelTrain datasetTESTOOD-GPT3OOD-REALTECGTEST-CC
Style Scalpel / stylometric modelTRAIN98.6817.6097.52100.014.67
Style Scalpel / stylometric modelTRAIN-CG98.2315.2097.5095.809.55
Style Scalpel / stylometric modelTRAIN + GPT-398.7998.9095.65100.023.18

Interpretation and limitations

The model is competitive on the main TEST split and strong on OOD-REAL and TECG under these conditions. OOD-GPT3 performance is weak when GPT-3 examples are absent from training, and TEST-CC remains weak. The project documentation notes that TEST-CC appears to contain substantial encoding or extraction artifacts.

Benchmark performance on specific datasets does not establish universal AI-detection accuracy. Results depend on domain, genre, model family, text length, editing, and other conditions.

These results are benchmark evidence, not a guarantee and not proof that a particular document has a particular origin.

Benchmark source

Abdalla, M. H. I., Malberg, S., Dementieva, D., Mosca, E., & Groh, G. (2023). A benchmark dataset to distinguish human-written and machine-generated scientific papers. Information, 14(10), 522. Read the benchmark paper.

View the complete comparison table in the project repository or read about the research behind Style Scalpel.