^start of line
Group frequency(?<frequency>[0-9]+) +matches at least one character from the set below:
0-9one character from 0 (4810) to 9 (5710)
\W+any character that is not a Unicode word character, repeated at least once
Group word(?<word>\pL+(?:,\h\pL+|\W)*) \pL+any Unicode letter, repeated at least once
Non-Capturing Group(?:,\h\pL+|\W)* *repeats the group contents below any number of times:
,the literal , (4410 / 548 / 2C16)
\hany horizontal whitespace character
\pL+any Unicode letter, repeated at least once
\Wany character that is not a Unicode word character
\h+any horizontal whitespace character, repeated at least once
Group gender(?<gender> [\pL()]+ (?:, \h* [\pL()]+)* ) +matches at least one character from the set below:
\pLany Unicode letter
(the literal ( (4010 / 508 / 2816)
)the literal ) (4110 / 518 / 2916)
Non-Capturing Group(?:, \h* [\pL()]+)* *repeats the group contents below any number of times:
\h*any horizontal whitespace character, repeated any number of times
+matches at least one character from the set below:
\pLany Unicode letter
(the literal ( (4010 / 508 / 2816)
)the literal ) (4110 / 518 / 2916)
\h+any horizontal whitespace character, repeated at least once
Group word_en(?<word_en> [^•]*[^•\s]) Negated Character Class[^•]* *matches any number of characters outside the set below:
•the literal • (822610 / 200428 / 202216)
Negated Character Class[^•\s] •the literal • (822610 / 200428 / 202216)
\sa Unicode whitespace character
\h*any horizontal whitespace character, repeated any number of times
\Ra newline sequence
\h*any horizontal whitespace character, repeated any number of times
Group sent_esp(?<sent_esp> [^–]*[^\s–] ) Negated Character Class[^–]* *matches any number of characters outside the set below:
–the literal – (821110 / 200238 / 201316)
Negated Character Class[^\s–] \sa Unicode whitespace character
–the literal – (821110 / 200238 / 201316)
\s*a Unicode whitespace character, repeated any number of times
–the literal – (821110 / 200238 / 201316)
\s*a Unicode whitespace character, repeated any number of times
Group sent_en(?<sent_en> .* (?:\R .*)*? ) .*any Unicode character (except for line terminators), repeated any number of times
Non-Capturing Group(?:\R .*)*? *?repeats the group contents below any number of times:
\Ra newline sequence
.*any Unicode character (except for line terminators), repeated any number of times
\h*any horizontal whitespace character, repeated any number of times
\Ra newline sequence
Group num1(?<num1> [0-9]+) +matches at least one character from the set below:
0-9one character from 0 (4810) to 9 (5710)
\h*any horizontal whitespace character, repeated any number of times
\|the literal | (12410 / 1748 / 7C16)
\h*any horizontal whitespace character, repeated any number of times
.*any Unicode character (except for line terminators), repeated any number of times
\Sany character that is not Unicode whitespace
g modifier: global. Finds all matches instead of stopping after the first
u modifier: unicode. Treats the pattern and subject as UTF-16
x modifier: extended. Ignores unescaped whitespace and comments outside character classes
m modifier: multiline. Causes ^ and $ to match the start and end of each line, not only the start and end of the string