Regular Expressions 101

Community Patterns

23

Get path from any text

Created·2023-01-31 14:38
Updated·2023-07-23 20:17
Flavor·PCRE2 (PHP)
Recommended·
Get path (windows style) from any type of text (error message, e-mail corps ...), quoted or not. THIS IS THE SINGLE LINE VERSION ! If you want understand how it work or edit it, go https://regex101.com/r/7o2fyy Relative path are not supported The goal is to catch what "Look like" a path. See the limitations UNC path and prefix path like //./], [//?/] or [//./UNC/] are allowed some url path like [file:///C:/] or [file://] are allowed Catch path quoted with ["] and [']. But these quotes are include with the catch Quoted path is not concerned by limitations Limitations : (only unquoted path) [dot] and [space] is allowed, but not in a row [dot+space] or [space+dot at end of file name isn't catched INSIDE A NAME FILE (or last directory if it is a path to a directory) : [comma] is not supported (it stop the catch) after a first [dot], any [space] stop the catch after a [space], catch is stoped if next character is not a [letter], [digit] or [-] so, double [space] stop the catch Compatibility compatible PCRE, PCRE2 AutoHotkey : don't forget to escape "%" in "`%" /!\ Powershell and .Net /!\\ : this regex need some modification to be interpreted by powershell. You have to replace each (?&CapturGroupName) by \k. Use this powershell code to do this replacement : ` $powershellRegex = @' [Put here the regex to replace (?&CapturGroupName) with \k] '@ -replace '\(\?&(\w+)\)', '\k' ` This example code must return : [Put here the regex to replace \k with \k]
Submitted by nitrateag

Community Library Entry

1

Regular Expression
Created·2026-06-30 06:47
Updated·2026-07-02 06:44
Flavor·PCRE2 (PHP)

/
(?<![A-Za-z0-9]) # left boundary: don't start in the middle of a letter/number run (?![^<>]*>) # HTML guard: skip if a ">" is reachable before a "<" (i.e. inside a tag) (?<SMILES> # ================ capture: the whole SMILES ================ # ---- FIRST COMPONENT ---- (?: # an atom: \[[0-9]* # bracket atom: "[" , optional isotope digits, (?: (?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U) # a valid element symbol, |(?:se|as|[bcnops]) # or a lowercase aromatic atom, |\* # or a wildcard ) [^\]]*\] # then charge / H-count / chirality, up to "]" |(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*) # OR an unbracketed organic-subset atom ) (?: # optional chain after the atom: (?: # zero+ interior tokens... (?:\[[0-9]*(?:(?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U)|(?:se|as|[bcnops])|\*)[^\]]*\]|(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*)) # an atom (same shape as above), |[=#$:\/\\\-] # or a bond ( = # $ : / \ - ), |[()] # or a branch paren, |(?:%[0-9]{2}|[0-9]) # or a ring closure ( %NN or single digit ) )* (?: # ...then a required terminal token (never a dangling bond): (?:\[[0-9]*(?:(?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U)|(?:se|as|[bcnops])|\*)[^\]]*\]|(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*)) # an atom, |%[0-9]{2} # or a two-digit ring closure, |[0-9)] # or a closing digit / paren ) )? (?: # ---- ADDITIONAL "."-SEPARATED COMPONENTS (salts, hydrates, ions) ---- \. # component separator "." (?:\[[0-9]*(?:(?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U)|(?:se|as|[bcnops])|\*)[^\]]*\]|(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*)) # first atom of this component (?: # optional chain after the atom: (?: # zero+ interior tokens... (?:\[[0-9]*(?:(?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U)|(?:se|as|[bcnops])|\*)[^\]]*\]|(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*)) # an atom (same shape as above), |[=#$:\/\\\-] # or a bond ( = # $ : / \ - ), |[()] # or a branch paren, |(?:%[0-9]{2}|[0-9]) # or a ring closure ( %NN or single digit ) )* (?: # ...then a required terminal token (never a dangling bond): (?:\[[0-9]*(?:(?:H[eogsf]?|L[iavru]|B[eahkri]?|C[arofmusenld]?|N[eiahopdb]?|O[sg]?|F[rle]?|M[godtcn]|A[lrsgutmc]|S[icerngmb]?|P[uabotmrd]?|Kr?|T[icebmsalh]|V|Z[nr]|G[ade]|R[buhenagf]|Yb?|I[nr]?|Xe|E[urs]|D[ysb]|W|U)|(?:se|as|[bcnops])|\*)[^\]]*\]|(?:Cl|Br|[BCNOFPSI]|[bcnops]|\*)) # an atom, |%[0-9]{2} # or a two-digit ring closure, |[0-9)] # or a closing digit / paren ) )? )* ) # ================ end capture ================ (?![A-Za-z0-9]) # right boundary: don't end in the middle of a letter/number run
/
gmx
Open regex in editor

Description

This matches some pretty complicated SMILES structures, works well for my app.

Submitted by Justin Hyland