Home/AI Detector Technologies/Accuracy & Results

Accuracy & Results

Performance metrics for the GetSolved AI Detector across AI model classes, text lengths, and content types. All figures are measured on held-out evaluation sets not used during training.

Overall Performance

On a balanced evaluation set (equal parts human and AI-generated text, 200–800 words each), the detector achieves approximately 92% overall accuracy. Performance varies by AI model and text length.

Approximately 8% of human texts may be incorrectly flagged as AI-generated, and approximately 6% of AI texts may go undetected. These rates are higher for very short texts and for content that falls in the borderline zone where human and AI writing characteristics overlap.

These figures represent aggregate performance across all supported model classes. Per-model accuracy is shown in the breakdown below.
Overall Accuracy
92%
balanced eval set
Human flagged as AI
~8%
false positive rate
AI text missed
~6%
false negative rate
Confidence Spectrum
Human
Borderline
AI
0%30%70%100%

Accuracy by AI Model

GPT-generated text is detected most reliably because its writing patterns are the most distinctive. Claude and Gemini follow closely. Text from less common or newer AI models may have lower detection accuracy because the classifier has less training data for these sources.

Human text is correctly classified in approximately 92% of cases. The remaining 8% typically falls in the borderline zone where writing characteristics overlap with heavily-edited or highly formulaic text.

When results fall in the borderline zone, we recommend reviewing the sentence-level heatmap and per-model confidence breakdown for a more complete picture.
detection_accuracy
GPT (3.5 / 4)96%
Claude94%
Gemini91%
Other AI models83%
Human text92%

Effect of Text Length

Accuracy scales with text length. The detector needs enough text to establish reliable patterns. For very short texts (under 50 words), results are unreliable and should be treated with caution.

The recommended range is 300–2,000 words. For documents longer than 2,000 words, the detector automatically segments the text and analyzes each section independently, which keeps accuracy high while identifying mixed-authorship regions.

accuracy_by_length
LengthAccuracyStatus
< 50 words74%Not recommended
50 – 100 words85%Low confidence
100 – 300 words91%Good
300 – 800 words94%Recommended
800 – 2,000 words96%Best results
> 2,000 words95%Segmented analysis

Understanding Your Results

Every scan produces several data points. Understanding what each one means helps you make informed decisions rather than relying on a single number.

The overall AI probability is the headline number, but the supporting data — model attribution, sentence heatmap, and segment analysis — often tells a richer story. A 55% AI probability with a strong GPT attribution and clear heatmap clusters is more informative than the percentage alone suggests.

Results should always be considered alongside context. The detector is a tool to assist human judgment, not replace it.
AI Probability (0–100%)

The overall likelihood that the text was generated by AI. Values below 30% suggest human authorship. Values above 70% strongly suggest AI. The 30–70% range is the borderline zone where additional review is recommended.

Example: 82% AI probability means the detector is fairly confident the text was AI-generated.

Model Attribution

Which AI model most likely produced the text, with per-model confidence percentages. This does not mean other models could not have generated it — it indicates the best statistical match.

Example: GPT 74%, Claude 18%, Gemini 8% means the text most closely resembles GPT output.

Sentence Heatmap

A color-coded view of AI probability for each individual sentence. Green sentences are likely human; red sentences are likely AI. This is especially useful for identifying partially AI-assisted documents.

Example: A green introduction followed by red body paragraphs suggests the author used AI for the main content.

Segment Analysis

For longer documents, the detector automatically splits the text at points where the writing style changes significantly. Each segment gets its own AI probability and model attribution.

Example: Segment 1 (paragraphs 1–2): 12% AI. Segment 2 (paragraphs 3–5): 89% AI (GPT).

Best Practices

Follow these guidelines to get the most reliable results from the AI Detector. The detector performs best when given enough text in a supported format.

These recommendations are based on our evaluation data and common patterns we see in real-world usage.

Submit enough text

Aim for at least 300 words. The 300–2,000 word range produces the most reliable results. Avoid submitting single sentences or very short paragraphs.

Use plain text

Remove formatting, headers, footers, and references before scanning. HTML tags, bullet-point lists, and tabular data can interfere with the linguistic analysis.

Scan the full document

For longer documents, submit the full text rather than selected passages. The detector's automatic segmentation will handle mixed-authorship detection more accurately.

Review borderline results

If the AI probability falls between 30% and 70%, check the sentence heatmap and model attribution. Use the result as one input alongside your own judgment.

Check the heatmap

Even when the overall probability is moderate, the sentence-level heatmap can reveal specific passages that were likely AI-generated — useful for targeted review.

Download the report

For formal or academic use, download the PDF AI detection report. It includes all results, timestamps, and the detector version — suitable for documentation purposes.

Known Limitations

No AI detector is 100% accurate. Understanding where accuracy drops helps set appropriate expectations and guides how results should be interpreted.

The detector is optimized for English-language text. Other languages are not supported and will produce unreliable results.

Short texts

Texts under 50 words may produce unreliable results. We recommend submitting at least 100 words for a meaningful analysis.

Heavily edited AI text

When AI output is substantially rewritten by a human, the detection confidence may decrease. The sentence-level heatmap can still highlight suspicious sections.

Non-English text

The detector is calibrated exclusively for English. Submitting text in other languages will produce inaccurate results.

Newer AI models

Models released after the latest training update may not be attributed correctly. The engine is updated regularly to incorporate new AI models.