Text tools

Normalize Persian and Arabic text, repair half-spaces, detect language and direction, and build URL-safe slugs. Each tool is described with its exact behavior and edge cases.

On this page
  1. Normalizer
    1. normalize()
    2. fixHalfSpace()
    3. clean()
  2. Detector
    1. isPersian() and isArabic()
    2. Direction and RTL
    3. isRtlLocale()
  3. Slugify
    1. The separator

Normalizer#

RtlyKit\Text\Normalizer makes Persian text comparable and storable. User input mixes Arabic and Persian letter shapes (an Arabic keyboard types «ي» and «ك»), vowel marks, tatweel and stray whitespace. The same word then looks identical on screen but fails a string comparison or a database search.

normalize()#

Normalizer::normalize($text, $removeDiacritics = true) does the following:

  • Arabic letters become their Persian counterparts: «ك» to «ک», «ي» «ى» «ے» to «ی», «ة» «ۀ» to «ه», «ؤ» to «و», «إ» «أ» «ٱ» to «ا».
  • Arabic-Indic digits (٠-٩) become Persian digits (۰-۹). English digits stay.
  • Tatweel (the stretching character U+0640) is removed.
  • Harakat, the short-vowel marks (U+064B to U+065F and U+0670), are removed. Pass false as the second argument to keep them.
  • Runs of whitespace (spaces, tabs, newlines) become a single space, and the ends are trimmed.
<?php
require 'vendor/autoload.php';

use RtlyKit\Text\Normalizer;
use function RtlyKit\normalize_text;

echo Normalizer::normalize("علي  كتاب  ١٢٣ 456"), "\n";    // علی کتاب ۱۲۳ 456
echo Normalizer::normalize("كــتاب"), "\n";                 // کتاب
echo Normalizer::normalize("  سلام \n\t دنیا  "), "\n";     // سلام دنیا
echo Normalizer::normalize("أحمد إبراهيم مؤمن ة"), "\n";    // احمد ابراهیم مومن ه
echo Normalizer::normalize("ثَبْت  ١٢٣ 456"), "\n";         // ثبت ۱۲۳ 456
echo Normalizer::normalize("ثَبْت", false), "\n";           // ثَبْت (diacritics kept)
echo normalize_text('كتاب'), "\n";                          // کتاب
What it leaves alone. The standalone hamza «ء» and «ئ» are not touched, because they carry meaning in Persian and Arabic words. The half-space (ZWNJ) is kept: «می‌خواهم» comes out byte for byte the same. English letters are unchanged.

The helper normalize_text() calls Normalizer::normalize() with the defaults.

fixHalfSpace()#

The half-space (ZWNJ, U+200C) is what makes «می‌خواهم» look right. Copy and paste often leaves several in a row, or one next to a space or at the end of the text. fixHalfSpace() cleans that up:

  • Soft hyphens (U+00AD) are removed.
  • A run of ZWNJs becomes one.
  • A ZWNJ right before or after whitespace is removed.
  • ZWNJs at the very start and end of the text are removed.
$a = "می\u{200C}\u{200C}\u{200C}خواهم";
echo preg_match_all('/./u', $a), ' ', preg_match_all('/./u', Normalizer::fixHalfSpace($a)), "\n";   // 10 8

$b = "\u{200C}سلام \u{200C}دنیا\u{200C}";
echo preg_match_all('/./u', $b), ' ', preg_match_all('/./u', Normalizer::fixHalfSpace($b)), "\n";   // 12 9
echo Normalizer::fixHalfSpace($b), "\n";                                  // سلام دنیا

It does not add missing half-spaces and does not merge ordinary spaces.

clean()#

Normalizer::clean() is the all-in-one version for storage and search. It runs normalize(), then fixHalfSpace(), then removes the invisible characters zero-width space (U+200B), zero-width joiner (U+200D) and the byte order mark (U+FEFF). ZWNJ is the only zero-width character it keeps.

$c = "كتاب\u{200B}ي  ١٢٣";
echo Normalizer::clean($c), "\n";   // کتابی ۱۲۳

normalize() alone would leave the invisible zero-width space inside the word. That is why clean() exists. Use it before you store or index user text.

Detector#

RtlyKit\Text\Detector answers questions about script, language and direction. All methods are static and never throw.

isPersian() and isArabic()#

Both are heuristics for Arabic-script text. Persian is recognised by letters that exist only in Persian (پ چ ژ گ, the Persian forms ک and ی, and Persian digits). Arabic is recognised by ي ك ى ة and Arabic-Indic digits. Letters that both languages share, including «ه» and the hamza forms, are not counted. The text counts as Persian when Persian-only letters are at least as frequent as Arabic-only letters. So ties, and text with no telling letters, count as Persian.

use RtlyKit\Text\Detector;

var_dump(Detector::isPersian('سلام'));        // bool(true), no telling letters, Persian by default
var_dump(Detector::isPersian('گچپژ'));        // bool(true)
var_dump(Detector::isArabic('مرحبا بكم'));    // bool(true)
var_dump(Detector::isArabic('كتاب ي'));       // bool(true)
var_dump(Detector::isPersian('hello'));       // bool(false), no Arabic script at all
var_dump(Detector::isPersian('۱۲۳'));         // bool(true), Persian digits
var_dump(Detector::isArabic('٣٤٥'));          // bool(true), Arabic-Indic digits
var_dump(Detector::isPersian(''));            // bool(false)
Good to know. This is a letter-count heuristic and not a language identifier. Very short text, names and Arabic words written in Persian can be ambiguous, and text that mixes both can go either way. Use it to choose a sensible default, not for decisions that must be exact.

Direction and RTL#

  • containsRtl($text): whether the text has any character from a right-to-left script (Hebrew, Arabic and its supplements, Syriac, Thaana, N'Ko and others, Arabic presentation forms) or the RLM mark.
  • direction($text): 'rtl' or 'ltr', taken from the first strong letter, following the Unicode bidirectional rule. Digits, punctuation and spaces are neutral. Text with no letters returns 'ltr'.
  • isHebrew($text): whether the text has Hebrew letters.
var_dump(Detector::containsRtl('abc'));           // bool(false)
echo Detector::direction('سلام hello'), "\n";    // rtl
echo Detector::direction('Hello مرحبا'), "\n";   // ltr (the first letter is Latin)
echo Detector::direction('۱۲۳'), "\n";           // ltr (digits are neutral, there is no letter)
var_dump(Detector::isHebrew('שלום'));            // bool(true)

The helpers contains_rtl() and text_direction() wrap containsRtl() and direction(). Use direction() for an HTML dir attribute on user-written text.

isRtlLocale()#

isRtlLocale($locale) tells you if a locale tag is written right to left. It splits on -, _, . and @, ignores case and surrounding spaces, and returns true when the language is a right-to-left one (fa, ar, he, ur, ps, ug, dv, ckb, azb and others) or when a script subtag such as Arab or Hebr is present.

var_dump(Detector::isRtlLocale('fa_IR'));     // bool(true)
var_dump(Detector::isRtlLocale('ar-SA'));     // bool(true)
var_dump(Detector::isRtlLocale('ckb-IQ'));    // bool(true)
var_dump(Detector::isRtlLocale('az-Arab'));   // bool(true), the script subtag decides
var_dump(Detector::isRtlLocale('ku'));        // bool(false), plain "ku" is not in the RTL language list
var_dump(Detector::isRtlLocale('en'));        // bool(false)
var_dump(Detector::isRtlLocale(''));          // bool(false)

Slugify#

Slugify::make($text, $separator = '-') turns a title into a URL-friendly slug and keeps Persian and Arabic letters readable. It does not transliterate.

  1. The text is cleaned with Normalizer::normalize().
  2. Whitespace and half-spaces become the separator.
  3. Everything except letters (any script), combining marks, digits (any set), hyphens and underscores is dropped. Punctuation and emoji disappear.
  4. Repeated hyphens, underscores and separators become one, and separators are trimmed from both ends.
  5. The result is lower-cased as UTF-8.
use RtlyKit\Text\Slugify;

echo Slugify::make('سلام دنیا'), "\n";               // سلام-دنیا
echo Slugify::make("کتاب\u{200C}خانه ملی"), "\n";    // کتاب-خانه-ملی
echo Slugify::make('Hello, World!'), "\n";           // hello-world
echo Slugify::make('  A--B__C  '), "\n";             // a-b_c
echo Slugify::make('سلام!!! دنیا؟'), "\n";           // سلام-دنیا
echo Slugify::make('۱۲۳ ابر'), "\n";                 // ۱۲۳-ابر
echo Slugify::make('كتاب ي'), "\n";                  // کتاب-ی
echo Slugify::make('سلام 😀 دنیا'), "\n";            // سلام-دنیا
echo Slugify::make('عکس.jpg'), "\n";                 // عکسjpg
var_dump(Slugify::make('$%^&'));                    // string(0) ""

Two things to keep in mind. Digits keep the set they were typed in (۱۲۳ stays Persian), so run Digits::toEnglish() first if you want ASCII digits in URLs. A dot is dropped and not treated as a separator, so a file extension joins the name (عکسjpg). Slugify the name and the extension separately.

The separator#

Pass any string as the second argument: '_', a longer string, or an empty string to join the words with nothing between them.

echo Slugify::make('سلام دنیا', '_'), "\n";        // سلام_دنیا
echo Slugify::make('سلام دنیا', ''), "\n";         // سلامدنیا
echo Slugify::make('Hello World', '--'), "\n";     // hello--world

The separator is limited to 64 bytes and must be valid UTF-8. A longer one throws RtlyKitException with the code input_too_long. An invalid one throws it with invalid_argument:

try {
    Slugify::make('Hello World', str_repeat('x', 65));
} catch (\RtlyKit\Exceptions\RtlyKitException $e) {
    echo $e->getMessage(), ' [', $e->getErrorCode()->value, "]\n";
    // The slug separator is too long. [input_too_long]
}
Good to know. A title made only of punctuation gives an empty string, and different titles can give the same slug. Handle the empty case yourself (for example, fall back to an id) and enforce uniqueness in your storage layer.
Encoding. Slugify needs valid UTF-8 for both the text and the separator. Invalid bytes in the text throw RtlyKitException with the code invalid_argument, because a slug is an identifier and bad bytes should not pass through. Normalizer is more forgiving. It works on a best-effort basis and returns input it cannot process as it was. Lower-casing uses the Unicode simple case mapping and needs no PHP extension.