docs: scope the foreign-script and surefire claims to what the code does

hasForeignScript detects by alphabet, so Latin-script leakage such as
the observed French "contiennent" is not caught. Say so in the javadoc
and admit the gap in the design doc rather than implying coverage.

failIfNoTests catches a misplaced or misnamed test class, not a
disabled one: an @Disabled class still reports as skipped and the
build stays green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F3jSvSTrG4qsniLSr6TpC6
This commit is contained in:
marcos
2026-08-05 15:09:37 +00:00
parent 4f427659bb
commit 3d2150d2a6
3 changed files with 21 additions and 5 deletions
@@ -143,11 +143,17 @@ no function that could do otherwise.
| failure | behaviour | | failure | behaviour |
| --- | --- | | --- | --- |
| empty content | retry once at a higher token ceiling, then apologise | | empty content | retry once at a higher token ceiling, then apologise |
| reply contains CJK or other foreign script | discard, retry once | | reply contains a non-Latin script (CJK, Cyrillic, Arabic, …) | discard, retry once |
| wiki 403 / timeout / no hit | answer without the article, and say the wiki was not consulted | | wiki 403 / timeout / no hit | answer without the article, and say the wiki was not consulted |
| MiniMax non-zero `base_resp` | log and apologise; HTTP 200 does not mean success | | MiniMax non-zero `base_resp` | log and apologise; HTTP 200 does not mean success |
| tool call absent despite forcing | fall back to answering ungrounded | | tool call absent despite forcing | fall back to answering ungrounded |
Known gap: the foreign-script check works by alphabet, so it only catches
non-Latin scripts. Latin-script leakage — the observed French `contiennent` in
an otherwise Portuguese answer — passes straight through, and this design does
not close that. Catching it would need dictionary or language-identification
work that is out of scope here.
## Features ## Features
- **Privacy per question.** `/ia` public, `/iap` visible only to the asker. - **Privacy per question.** `/ia` public, `/iap` visible only to the asker.
+3 -2
View File
@@ -71,8 +71,9 @@
<configuration> <configuration>
<!-- The pin above guards against a surefire version that quietly stops <!-- The pin above guards against a surefire version that quietly stops
discovering tests. It does not guard against a test class being discovering tests. It does not guard against a test class being
misplaced, misnamed or disabled, which ends the same way: a green misplaced or misnamed, which ends the same way: a green build that
build that ran nothing. Failing on an empty suite closes that gap. --> ran nothing. Failing on an empty suite closes that gap. A disabled
class is still reported as skipped, so it is not covered here. -->
<failIfNoTests>true</failIfNoTests> <failIfNoTests>true</failIfNoTests>
</configuration> </configuration>
</plugin> </plugin>
@@ -11,8 +11,8 @@ import java.util.regex.Pattern;
final class AiText { final class AiText {
/** /**
* Scripts that should never appear in a Portuguese answer. The model has * Non-Latin scripts that should never appear in a Portuguese answer. The
* been observed dropping single Chinese words mid-sentence. * model has been observed dropping single Chinese words mid-sentence.
*/ */
private static final Pattern FOREIGN = Pattern.compile( private static final Pattern FOREIGN = Pattern.compile(
"[\\p{IsHan}\\p{IsHiragana}\\p{IsKatakana}\\p{IsHangul}\\p{IsCyrillic}\\p{IsArabic}]"); "[\\p{IsHan}\\p{IsHiragana}\\p{IsKatakana}\\p{IsHangul}\\p{IsCyrillic}\\p{IsArabic}]");
@@ -20,6 +20,15 @@ final class AiText {
private AiText() { private AiText() {
} }
/**
* True if the text contains a character from a non-Latin script.
*
* <p>This detects leakage by alphabet, so it catches only what a different
* alphabet makes visible. A foreign word written in Latin script is
* <b>not</b> caught: the French {@code contiennent}, observed in an
* otherwise Portuguese reply, passes this check. Catching that would need
* dictionary or language-identification work this method does not do.
*/
static boolean hasForeignScript(String text) { static boolean hasForeignScript(String text) {
return text != null && FOREIGN.matcher(text).find(); return text != null && FOREIGN.matcher(text).find();
} }