On this page
Normalizer#
RtlyKit\Text\Normalizer makes Persian text comparable and storable. User input mixes Arabic and Persian letter shapes (an Arabic keyboard types «ي» and «ك»), vowel marks, tatweel and stray whitespace. The same word then looks identical on screen but fails a string comparison or a database search.
normalize()#
Normalizer::normalize($text, $removeDiacritics = true) does the following:
- Arabic letters become their Persian counterparts: «ك» to «ک», «ي» «ى» «ے» to «ی», «ة» «ۀ» to «ه», «ؤ» to «و», «إ» «أ» «ٱ» to «ا».
- Arabic-Indic digits (
٠-٩) become Persian digits (۰-۹). English digits stay. - Tatweel (the stretching character U+0640) is removed.
- Harakat, the short-vowel marks (U+064B to U+065F and U+0670), are removed. Pass
falseas the second argument to keep them. - Runs of whitespace (spaces, tabs, newlines) become a single space, and the ends are trimmed.
<?php
require 'vendor/autoload.php';
use RtlyKit\Text\Normalizer;
use function RtlyKit\normalize_text;
echo Normalizer::normalize("علي كتاب ١٢٣ 456"), "\n"; // علی کتاب ۱۲۳ 456
echo Normalizer::normalize("كــتاب"), "\n"; // کتاب
echo Normalizer::normalize(" سلام \n\t دنیا "), "\n"; // سلام دنیا
echo Normalizer::normalize("أحمد إبراهيم مؤمن ة"), "\n"; // احمد ابراهیم مومن ه
echo Normalizer::normalize("ثَبْت ١٢٣ 456"), "\n"; // ثبت ۱۲۳ 456
echo Normalizer::normalize("ثَبْت", false), "\n"; // ثَبْت (diacritics kept)
echo normalize_text('كتاب'), "\n"; // کتاب
The helper normalize_text() calls Normalizer::normalize() with the defaults.
fixHalfSpace()#
The half-space (ZWNJ, U+200C) is what makes «میخواهم» look right. Copy and paste often leaves several in a row, or one next to a space or at the end of the text. fixHalfSpace() cleans that up:
- Soft hyphens (U+00AD) are removed.
- A run of ZWNJs becomes one.
- A ZWNJ right before or after whitespace is removed.
- ZWNJs at the very start and end of the text are removed.
$a = "می\u{200C}\u{200C}\u{200C}خواهم";
echo preg_match_all('/./u', $a), ' ', preg_match_all('/./u', Normalizer::fixHalfSpace($a)), "\n"; // 10 8
$b = "\u{200C}سلام \u{200C}دنیا\u{200C}";
echo preg_match_all('/./u', $b), ' ', preg_match_all('/./u', Normalizer::fixHalfSpace($b)), "\n"; // 12 9
echo Normalizer::fixHalfSpace($b), "\n"; // سلام دنیا
It does not add missing half-spaces and does not merge ordinary spaces.
clean()#
Normalizer::clean() is the all-in-one version for storage and search. It runs normalize(), then fixHalfSpace(), then removes the invisible characters zero-width space (U+200B), zero-width joiner (U+200D) and the byte order mark (U+FEFF). ZWNJ is the only zero-width character it keeps.
$c = "كتاب\u{200B}ي ١٢٣";
echo Normalizer::clean($c), "\n"; // کتابی ۱۲۳
normalize() alone would leave the invisible zero-width space inside the word. That is why clean() exists. Use it before you store or index user text.
Detector#
RtlyKit\Text\Detector answers questions about script, language and direction. All methods are static and never throw.
isPersian() and isArabic()#
Both are heuristics for Arabic-script text. Persian is recognised by letters that exist only in Persian (پ چ ژ گ, the Persian forms ک and ی, and Persian digits). Arabic is recognised by ي ك ى ة and Arabic-Indic digits. Letters that both languages share, including «ه» and the hamza forms, are not counted. The text counts as Persian when Persian-only letters are at least as frequent as Arabic-only letters. So ties, and text with no telling letters, count as Persian.
use RtlyKit\Text\Detector;
var_dump(Detector::isPersian('سلام')); // bool(true), no telling letters, Persian by default
var_dump(Detector::isPersian('گچپژ')); // bool(true)
var_dump(Detector::isArabic('مرحبا بكم')); // bool(true)
var_dump(Detector::isArabic('كتاب ي')); // bool(true)
var_dump(Detector::isPersian('hello')); // bool(false), no Arabic script at all
var_dump(Detector::isPersian('۱۲۳')); // bool(true), Persian digits
var_dump(Detector::isArabic('٣٤٥')); // bool(true), Arabic-Indic digits
var_dump(Detector::isPersian('')); // bool(false)
Direction and RTL#
containsRtl($text): whether the text has any character from a right-to-left script (Hebrew, Arabic and its supplements, Syriac, Thaana, N'Ko and others, Arabic presentation forms) or the RLM mark.direction($text):'rtl'or'ltr', taken from the first strong letter, following the Unicode bidirectional rule. Digits, punctuation and spaces are neutral. Text with no letters returns'ltr'.isHebrew($text): whether the text has Hebrew letters.
var_dump(Detector::containsRtl('abc')); // bool(false)
echo Detector::direction('سلام hello'), "\n"; // rtl
echo Detector::direction('Hello مرحبا'), "\n"; // ltr (the first letter is Latin)
echo Detector::direction('۱۲۳'), "\n"; // ltr (digits are neutral, there is no letter)
var_dump(Detector::isHebrew('שלום')); // bool(true)
The helpers contains_rtl() and text_direction() wrap containsRtl() and direction(). Use direction() for an HTML dir attribute on user-written text.
isRtlLocale()#
isRtlLocale($locale) tells you if a locale tag is written right to left. It splits on -, _, . and @, ignores case and surrounding spaces, and returns true when the language is a right-to-left one (fa, ar, he, ur, ps, ug, dv, ckb, azb and others) or when a script subtag such as Arab or Hebr is present.
var_dump(Detector::isRtlLocale('fa_IR')); // bool(true)
var_dump(Detector::isRtlLocale('ar-SA')); // bool(true)
var_dump(Detector::isRtlLocale('ckb-IQ')); // bool(true)
var_dump(Detector::isRtlLocale('az-Arab')); // bool(true), the script subtag decides
var_dump(Detector::isRtlLocale('ku')); // bool(false), plain "ku" is not in the RTL language list
var_dump(Detector::isRtlLocale('en')); // bool(false)
var_dump(Detector::isRtlLocale('')); // bool(false)
Slugify#
Slugify::make($text, $separator = '-') turns a title into a URL-friendly slug and keeps Persian and Arabic letters readable. It does not transliterate.
- The text is cleaned with
Normalizer::normalize(). - Whitespace and half-spaces become the separator.
- Everything except letters (any script), combining marks, digits (any set), hyphens and underscores is dropped. Punctuation and emoji disappear.
- Repeated hyphens, underscores and separators become one, and separators are trimmed from both ends.
- The result is lower-cased as UTF-8.
use RtlyKit\Text\Slugify;
echo Slugify::make('سلام دنیا'), "\n"; // سلام-دنیا
echo Slugify::make("کتاب\u{200C}خانه ملی"), "\n"; // کتاب-خانه-ملی
echo Slugify::make('Hello, World!'), "\n"; // hello-world
echo Slugify::make(' A--B__C '), "\n"; // a-b_c
echo Slugify::make('سلام!!! دنیا؟'), "\n"; // سلام-دنیا
echo Slugify::make('۱۲۳ ابر'), "\n"; // ۱۲۳-ابر
echo Slugify::make('كتاب ي'), "\n"; // کتاب-ی
echo Slugify::make('سلام 😀 دنیا'), "\n"; // سلام-دنیا
echo Slugify::make('عکس.jpg'), "\n"; // عکسjpg
var_dump(Slugify::make('$%^&')); // string(0) ""
Two things to keep in mind. Digits keep the set they were typed in (۱۲۳ stays Persian), so run Digits::toEnglish() first if you want ASCII digits in URLs. A dot is dropped and not treated as a separator, so a file extension joins the name (عکسjpg). Slugify the name and the extension separately.
The separator#
Pass any string as the second argument: '_', a longer string, or an empty string to join the words with nothing between them.
echo Slugify::make('سلام دنیا', '_'), "\n"; // سلام_دنیا
echo Slugify::make('سلام دنیا', ''), "\n"; // سلامدنیا
echo Slugify::make('Hello World', '--'), "\n"; // hello--world
The separator is limited to 64 bytes and must be valid UTF-8. A longer one throws RtlyKitException with the code input_too_long. An invalid one throws it with invalid_argument:
try {
Slugify::make('Hello World', str_repeat('x', 65));
} catch (\RtlyKit\Exceptions\RtlyKitException $e) {
echo $e->getMessage(), ' [', $e->getErrorCode()->value, "]\n";
// The slug separator is too long. [input_too_long]
}
RtlyKitException with the code invalid_argument, because a slug is an identifier and bad bytes should not pass through. Normalizer is more forgiving. It works on a best-effort basis and returns input it cannot process as it was. Lower-casing uses the Unicode simple case mapping and needs no PHP extension.