SlopShape: Structural Detection of AI-Generated Commercial Web Content
Key Result
SlopShape achieves 98 % macro‑F1 in detecting AI‑generated commercial blog posts using only structural features, and it attributes 79.3 % of AI posts to the correct model. This demonstrates that AI‑generated text leaves a consistent “shape” that survives aggressive rewording.
What the Study Tested
The authors replicated the StoryScope methodology (Russell et al., 2026) on commercial content. They assembled:
- 2,250 human‑written blog posts from 268 company domains (pre‑ChatGPT era).
- 11,250 AI‑generated mirrors created by five frontier language models, each post paired with a human counterpart.
- A 214‑feature instrument capturing structural signatures such as ordering of information, evidence placement, voice, and paragraph hierarchy.
- An LLM‑driven pipeline that extracts these features and feeds them to a classifier.
The instrument was validated with a human annotation gold set, achieving a human‑human Cohen’s κ of 0.928 and a human‑model κ of 0.946, indicating near‑perfect agreement on the structural labels.
Performance on Held‑Out Companies
When the classifier was evaluated on companies excluded from training, it retained 98.0 % macro‑F1. Re‑wording each AI post with its own generating model (i.e., full paraphrase) did not degrade performance (98.1 % macro‑F1). This confirms that the detection signal resides in the underlying structure rather than surface wording.
Attribution Capability
Beyond binary detection, SlopShape can attribute AI posts to the correct source model:
- Correct attribution rate: 79.3 %.
- Random baseline: 16.7 % (five models, uniform chance). The attribution signal stems from subtle differences in how each model organizes content, selects evidence, and adopts a narrative voice.
Human vs. AI Structural Profiles
The study reports that human‑written posts occupy rare structural configurations compared with AI‑generated ones, which tend to follow a tidy, self‑announcing shape. This aligns with StoryScope’s findings on fiction and suggests a broader, domain‑agnostic pattern.
Community Reactions
"Structure alone can differentiate AI content, interesting approach." – DylanMerigaud (HN comment)
"Immediately saw false positives on content written before ChatGPT. Not difficult to see how the methodology is wrong when it considers restating the thesis in the conclusion to be signal." – evantbyrne (HN comment)
"The idea is interesting but the tooling feels lossy; relying on LLM prompts for feature extraction may hinder reproducibility." – asdff (HN comment)
These comments highlight two recurring concerns:
- Potential false positives on older human content that incidentally matches AI‑style structures.
- Reliance on LLMs for feature extraction, which may introduce non‑deterministic behavior and dependence on external APIs.
Limitations and Open Questions
- False‑Positive Risk: Early‑era corporate blogs sometimes reuse formulaic structures (e.g., thesis‑statement‑conclusion) that the model may flag as AI.
- Determinism: The pipeline uses LLMs to parse HTML and generate feature vectors, raising reproducibility questions if the underlying model changes or is unavailable.
- Scope: The dataset focuses on commercial blog posts; applicability to other genres (news, academic writing) remains untested.
Resources Released
The authors provide a complete reproducibility package:
- Pipeline code (GitHub: https://github.com/pulse-energy-eu/slopshape)
- Feature instrument (214 structural descriptors)
- Prompt templates used for LLM‑based extraction
- Aggregated artifacts (trained classifiers, evaluation scripts)
Why It Matters
SlopShape shows that AI‑generated text leaves a persistent structural fingerprint detectable even after aggressive paraphrasing. This opens a new line of defense against AI‑generated misinformation and commercial spam, complementing traditional lexical detectors that fail under re‑writing. However, the approach’s dependence on LLM‑driven feature extraction and its susceptibility to false positives on formulaic human writing warrant further scrutiny before deployment in high‑stakes settings.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Project
- Dispatch