From: Snapshot-Content-Location: https://www3.ntu.edu.sg/home/ehchua/programming/howto/Regexe.html Subject: Regular Expression (Regex) Tutorial Date: Sun, 13 Aug 2023 12:24:56 -0000 MIME-Version: 1.0 Content-Type: multipart/related; type="text/html"; boundary="----MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD----" ------MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD---- Content-Type: text/html Content-ID: Content-Transfer-Encoding: quoted-printable Content-Location: https://www3.ntu.edu.sg/home/ehchua/programming/howto/Regexe.html Regular Expression (Regex) Tutorial

yet another insignificant pro= gramming notes...   |   HOME

TABLE OF CONTENTS (HIDE)
1.  Regex By Examples
1.1  Regex Syntax Summary1.2  Example: Numbers [0-9]+ or \= d+
1.3  Code Examples (P= ython, Java, JavaScript, Perl, PHP)
1.4&n= bsp; Example: Full Numeric Strings ^[0-9]+$ or ^\d+$1.5  Example: Positive Integer Litera= ls [1-9][0-9]*|0 or [1-9]\d*|0
1.6  Example: Full Integer Literals ^[+-]?[1-9][0-9]*|0$ or 1.7  = ;Example: Identifiers (or Names) [a-zA-= Z_][0-9a-zA-Z_]* or [a-zA-Z_]\w*=
1.8  Example: Image Fil= enames ^\w+\.(gif|png|jpg|jpeg)$=
1.9  Example: Email Addresses = ^\w+([.-]?\w+)*@\w+([.-]?\w+)*(\.\w{2,3= })+$
1.10  Example: Swa= pping Words using Parenthesized Back-References ^(\S+)\s+(\S+)$ and $2 = $1
1.11  Example: HTTP = Addresses ^http:\/\/\S+(\/\S+)*(\/)?$
1.12  Example: Regex Pat= terns in AngularJS
1.13  Examp= le: Sample Regex in Perl
2.  Reg= ular Expression (Regex) Syntax
2.1 &= nbsp;Matching a Single Character
2.2 = ; Regex Special Characters and Escape Sequences
2.3  Matching a Sequence of Characters (String or Tex= t)
2.4  OR (|) Operator
2.5  = Bracket List (Character Class) [...], [^...], [.-.]=
2.6  Metacharacters ., \w, \W, \d, \D, \s, \S
2.7  B= ackslash (\) and Regex Escape Sequences
2.8  Occurrence Indicators (Repet= ition Operators): +, *, ?, {m}, {m,n}, {m,}
2.9  Modifiers=
2.10  Greediness, Laziness an= d Backtracking for Repetition Operators
= 2.11  Position Anchors ^, $, \b, \B, \<, \>, \A, \Z
2.12&nbs= p; Capturing Matches via Parenthesized Back-References & Matched V= ariables $1, $2<= /span>, ...
2.13  (Advanced) L= ookahead/Lookbehind, Groupings and Conditional
2.14  Unicode
3. &nb= sp;Regex in Programming Languages

Regular Expressions (Regex)

Regular Expression, or regex or regexp in sho= rt, is extremely and amazingly powerful in searching and manipulating text = strings, particularly in processing text files. One line of regex can easil= y replace several dozen lines of programming codes.

Regex is supported in all the scripting languages (such as Perl, Python= , PHP, and JavaScript); as well as general purpose programming languages su= ch as Java; and even word processors such as Word for searching texts. Gett= ing started with regex may not be easy due to its geeky syntax, but it is c= ertainly worth the investment of your time.

1.  Regex By Examples

This section is meant for those who need to refresh their memory. For no= vices, go to the next section to learn the syntax, before looking at thes= e examples.

1.1  Regex Syntax Summary=

  • Character: All characters, except t= hose having special meaning in regex, matches themselves. E.g., the regex <= code>x matches substring "x"; regex 9 matc= hes "9"; regex =3D matches "=3D"; an= d regex @ matches "@".
  • Special Regex Characters: These cha= racters have special meaning in regex (to be discussed below): ., +, *, ?, ^, $= , (, ), [, ], {, }, |, \.
  • Escape Sequences (\char):
    • To match a character having special meaning in regex, you need to use a= escape sequence prefix with a backslash (\). E.g., \. matches "."; regex \+ matches "+"; and regex \( matches "(".
    • You also need to use regex \\ to match "\" (b= ack-slash).
    • Regex recognizes common escape sequences such as \n for ne= wline, \t for tab, \r for carriage-return, = \nnn for a up to 3-digit octal number, \xhh for a two-d= igit hex code, \uhhhh for a 4-digit Unicode, \uhhhhhhhh<= /code> for a 8-digit Unicode.
  • A Sequence of Characters (or String): Strings can be matched via combining a sequence of characters (called su= b-expressions). E.g., the regex Saturday matches "Saturd= ay". The matching, by default, is case-sensitive, but can be set to = case-insensitive via modifier.
  • OR Operator (|): E.g., the regex four|4 accepts strings "f= our" or "4".
  • Character class (or Bracket List):
    • [...]: Accept ANY ONE of the chara= cter within the square bracket, e.g., [aeiou] matches "a= ", "e", "i", "o" or "u"= .
    • [.-.] (Range Expression): Accept A= NY ONE of the character in the range, e.g., [0-9] mat= ches any digit; [A-Za-z] matches any uppercase or lowercase le= tters.
    • [^...]: NOT ONE of the character, = e.g., [^0-9] matches any non-digit.
    • Only these four characters require escape sequence inside the brack= et list: ^, -, ], \.
  • Occurrence Indicators (or Repetition Opera= tors):
    • +: one or more (1+), = e.g., [0-9]+ matches one or more digits such as '123', '000'.
    • *: zero or more (0+),= e.g., [0-9]* matches zero or more digits. It accepts all thos= e in [0-9]+ plus the empty string.
    • ?: zero or one (optional), e.g., <= code>[+-]? matches an optional "+", "-", o= r an empty string.
    • {m,n}: m to n (both inclusive)
    • {m}: exactly m times<= /li>
    • {m,}: m or more (m+)
  • Metacharacters: matches a character
    • . (dot): ANY ONE character except = newline. Same as [^\n]
    • \d, \D: ANY ONE digit/non-digit character. Digits are [0-9]
    • \w, \W: ANY ONE word/non-word character. For ASCII, word characters are [a-zA-Z0-9_]
    • \s, \S: ANY ONE space/non-space character. For ASCII, whitespace characters = are [ \n\r\t\f]
  • Position Anchors: does not match ch= aracter, but position such as start-of-line, end-of-line, start-of-word and= end-of-word.
    • ^, $: start-of-line and end-of-line respectively. E.g., ^[0-9]$= matches a numeric string.
    • \b: boundary of word, i.e., start-= of-word or end-of-word. E.g., \bcat\b matches the word "= cat" in the input string.
    • \B: Inverse of \b, i.e., non-start= -of-word or non-end-of-word.
    • \<, \= >: start-of-word and end-of-word respectively, similar to \= b. E.g., \<cat\> matches the word "cat" in the input string.
    • \A, \Z: start-of-input and end-of-input respectively.
  • Parenthesized Back References:
    • Use parentheses ( ) to create a back reference.
    • Use $1, $2, ... (Java, Perl, JavaScript) or <= code>\1, \2, ... (Python) to retreive the back referenc= es in sequential order.
  • Laziness (Curb Greediness for Repetition O= perators): *?, +?, ??, = {m,n}?, {m,}?

1.2  Example: Numbers [0-9]+ or \d+=

  1. A regex (regular expression) consists of a sequence of sub= -expressions. In this example, [0-9] and +.<= /li>
  2. The [...], known as character class (or brack= et list), encloses a list of characters. It matches any SINGLE charact= er in the list. In this example, [0-9] matches any SINGLE char= acter between 0 and 9 (i.e., a digit), where dash (-) denotes = the range.
  3. The +, known as occurrence indicator (or repe= tition operator), indicates one or more occurrences (1+) = of the previous sub-expression. In this case, [0-9]+ matches o= ne or more digits.
  4. A regex may match a portion of the input (i.e., substring) or the entir= e input. In fact, it could match zero or more substrings of the input (with= global modifier).
  5. This regex matches any numeric substring (of digits 0 to 9) of the inpu= t. For examples,
    1. If the input is "abc123xyz", it matches substring "1= 23".
    2. If the input is "abcxyz", it matches nothing.
    3. If the input is "abc00123xyz456_0", it matches substrings = "00123", "456" and "0" (three matche= s).
    Take note that this regex matches number with leading zeros, such as = "000", "0123" and "0001", which may not be= desirable.
  6. You can also write \d+, where \d is known as = a metacharacter that matches any digit (same as [0-9]= ). There are more than one ways to write a regex! Take note that many progr= amming languages (C, Java, JavaScript, Python) use backslash \= as the prefix for escape sequences (e.g., \n for newline), and you need to write "\\d+" instead.

1.3  Code Examples (Python, Java, JavaScript, Perl, PHP)

Code Example in Python

See "Python's re module for Regular = Expression" for full coverage.

Python supports Regex via module re. Python also uses backs= lash (\) for escape sequences (i.e., you need to write \= \ for \, \\d for \d), but it = supports raw string in the form of r'...', which ignore the in= terpretation of escape sequences - great for writing regex.

# Test under the=
 Python Command-Line Interpreter
$ python3
......
>>> import re   # N=
eed module 're' for regular expression

# Try find: re.findall(regexStr, inStr) -> matchedSubstringsList
# r'...' denotes raw strings which ignore escape code, i.e., r'\n' is '\'+'=
n'
>>> re.findall(r'[0-9]+', 'abc123xyz')
['123']   # Return a list of matched substrin=
gs
>>> re.findall(r'[0-9]+', 'abcxyz')
[]
>>> re.findall(r'[0-9]+', 'abc00123xyz456_0')
['00123', '456', '0']
>>> re.findall(r'\d+', 'abc00123xyz456_0')
['00123', '456', '0']

# Try substitute: re.sub(regexStr, <=
em>replacementStr, inStr) -> outStr
>>> re.sub(r'[0-9]+', r'*', 'abc00123xyz456_0')
'abc*xyz*_*'

# Try substitute with count: re.subn(rege=
xStr, replacementStr, inStr) -> (outStr,=
 count)
>>> re.subn(r'[0-9]+', r'*', 'abc00123xyz456_0')
('abc*xyz*_*', 3)   # Return a tuple of outpu=
t string and count
Code Example in Java

See "Regular Expressions (Regex) in Java" for full coverage.<= /p>

Java supports Regex in package java.util.regex.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
import java.util.regex.Pattern;
import java.util.regex.Matcher;

public class TestRegexNumbers {
   public static void main(String[] args) {

      String inputStr =3D "abc00123xyz456_0";  // Input String for matching
      String regexStr =3D "[0-9]+";            // Regex to be matched

      // Step 1: Compile a regex via static m=
ethod Pattern.compile(), default is case-sensitive
      Pattern pattern =3D Pattern.compile(regexSt=
r);
      // Pattern.compile(regex, Pattern.CASE_=
INSENSITIVE);  // for case-insensitive matching

      // Step 2: Allocate a matching engine f=
rom the compiled regex pattern,
      //         and bind to the input string=

      Matcher matcher =3D pattern.matcher(inputSt=
r);

      // Step 3: Perform matching and Process=
 the matching results
      // Try Matcher.find(), which finds the =
next match
      while (matcher.find()) {
         System.out.println("find() found substring \"" + matcher.group()
               + "\" starting at index " + matche=
r.start()
               + " and ending at index " + matche=
r.end());
      }

      // Try Matcher.matches(), which tries t=
o match the ENTIRE input (^...$)
      if (matcher.matches()) {
         System.out.println("matches() found substring \"" + matcher.group(=
)
               + "\" starting at index " + matcher.start()
               + " and ending at index " + matcher.end());
      } else {
         System.out.println("matches() found nothing");
      }

      // Try Matcher.lookingAt(), which tries=
 to match from the START of the input (^...)
      if (matcher.lookingAt()) {
         System.out.println("lookingAt() found substring \"" + matcher.grou=
p()
               + "\" starting at index " + matcher.start()
               + " and ending at index " + matcher.end());
      } else {
         System.out.println("lookingAt() found nothing");
      }

      // Try Matcher.replaceFirst(), which re=
places the first match
      String replacementStr =3D "**";
      String outputStr =3D matcher.replaceFirst(r=
eplacementStr); // first match only
      System.out.println(outputStr);

      // Try Matcher.replaceAll(), which repl=
aces all matches
      replacementStr =3D "++";
      outputStr =3D matcher.replaceAll(replacemen=
tStr); // all matches
      System.out.println(outputStr);
   }
}

The output is:

find() found substring "00123" starting at index 3 an=
d ending at index 8
find() found substring "456" starting at index 11 and ending at index 14
find() found substring "0" starting at index 15 and ending at index 16
matches() found nothing
lookingAt() found nothing
abc**xyz456_0
abc++xyz++_++
Code Example in Perl

See "Regular Expression (Regex) in Perl" for= full coverage.

Perl makes extensive use of regular expressions with many built-in synta= xes and operators. In Perl (and JavaScript), a regex is delimited by a pa= ir of forward slashes (default), in the form of /regex/. You can use built-in operators:

  • m/regex/modifier<= /span> or /regex/modifier<= /em>: Match against the regex. m = is optional.
  • s/regex/replacement/modifier: Substitute matched substring(s) by the replace= ment.

In Perl, you can use single-quoted non-interpolating string '....'= to write regex to disable interpretation of backslash (\) by Perl.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
#!/usr/bin/env perl
use strict;
use warnings;

my $inStr =3D 'abc00123xyz456_0';  # input st=
ring
my $regex =3D '[0-9]+';            # regex pa=
ttern string in non-interpolating string

# Try match /regex/modifiers (or m/regex/modi=
fiers)
my @matches =3D ($inStr =3D~ /$regex/g);  =
# Match $inStr with regex with global modifie=
r
                                      # Store=
 all matches in an array
print "@matches\n";   # Output: 00123 456 0

while ($inSt=
r =3D~ /$regex/g) {
   # The built-in array variables @- and @+ k=
eep the start and end positions
   #   of the matches, where $-[0] and $+[0] =
is the full match, and
   #   $-[n] and $+[n] for back references $1=
, $2, etc.
   print substr($inStr, $-[0], $+[0] - $-[0]), ', ';   # Output: 00123, 456, 0,
}
print "\n";

# Try substitute  s/regex/replacement/modifie=
rs
$inStr =3D~ s/$regex/**/g;    # with global modifier
print "$inStr\n";           # Output: abc**xy=
z**_**
Code Example in JavaScript

See "Regular Expression in JavaScript= " for full coverage.

In JavaScript (and Perl), a regex is delimited by a pair of forward sl= ashes, in the form of /.../. There are two sets of methods, is= sue via a RegEx object or a String object.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
<!DOCTYPE html>
<!-- JSRegexNumbers.html -->
<html lang=3D"en">
<head>
<meta charset=3D"utf-8">
<title>JavaScript Example: Regex</title>
<script>
var inStr =3D "abc123xyz456_7_00";

// Use RegExp.test(inStr) to check if inStr c=
ontains the pattern
console.log(/[0-9]+/.test(inStr));  // true

// Use String.search(regex) to check if the s=
tring contains the pattern
// Returns the start position of the matched =
substring or -1 if there is no match
console.log(inStr.search(/[0-9]+/));  // 3

// Use String.match() or RegExp.exec() to fin=
d the matched substring,
//   back references, and string index
console.log(inStr.match(/[0-9]+/));  // ["123=
", input:"abc123xyz456_7_00", index:3, length:"1"]
console.log(/[0-9]+/.exec(inStr));   // ["123=
", input:"abc123xyz456_7_00", index:3, length:"1"]

// With g (global) option
console.log(inStr.match(/[0-9]+/g));  // ["12=
3", "456", "7", "00", length:4]

// RegExp.exec() with g flag can be issued re=
peatedly.
// Search resumes after the last-found positi=
on (maintained in property RegExp.lastIndex).
var pattern =3D /[0-9]+/g;
var result;
while (result =3D pattern.exec(inStr)) {
   console.log(result);
   console.log(pattern.lastIndex);
      // ["123"],  6
      // ["456"], 12
      // ["7"],   14
      // ["00"],  17
}

// String.replace(regex, replacement):
console.log(inStr.replace(/\d+/, "**"));   //=
 abc**xyz456_7_00
console.log(inStr.replace(/\d+/g, "**"));  //=
 abc**xyz**_**_**
</script>
</head>
<body>
  <h1>Hello,</h1>
</body>
</html>
Code Example in PHP

[TODO]

1.4  Example: Full Numeric Strings ^[0-9]+$ or ^\d+$

  1. The leading ^ and the trailing $ are known as= position anchors, which match the start and end positions of the = line, respectively. As the result, the entire input string shall be matched= fully, instead of a portion of the input string (substring).
  2. This regex matches any non-empty numeric strings (comprising of digits = 0 to 9), e.g., "0" and "12345". It does not match= with "" (empty string), "abc", "a123", "ab= c123xyz", etc. However, it also matches "000", "0= 123" and "0001" with leading zeros.

1.5  Example: Positive Integer Litera= ls [1-9][0-9]*|0 or [1-9]\d*|0

  1. [1-9] matches any character between 1 to 9; [0-9]* matches zero or more digits. The * is an occurrence = indicator representing zero or more occurrences. Together, [1-9]= [0-9]* matches any numbers without a leading zero.
  2. | represents the OR operator; which is used to include the= number 0.
  3. This expression matches "0" and "123"; but do= es not match "000" and "0123" (but see below).
  4. You can replace [0-9] by metacharacter \d, bu= t not [1-9].
  5. We did not use position anchors ^ and $ in this regex. Hence, it can match any parts of the input string. For e= xamples,
    1. If the input string is "abc123xyz", it matches the substri= ng "123".
    2. If the input string is "abcxyz", it matches nothing.
    3. If the input string is "abc123xyz456_0", it matches substr= ings "123", "456" and "0" (three mat= ches).
    4. If the input string is "0012300", it matches substrings: <= code>"0", "0" and "12300" (three matches)!= !!

1.6  Example: Full Integer Literals ^[+-]?[1-9][0-9]*|0$ or ^[+-]?[1-9]\d*|0$

  1. This regex match an Integer literal (for entire string with the pos= ition anchors), both positive, negative and zero.
  2. [+-] matches either + or - sign= . ? is an occurrence indicator denoting 0 or 1 occurr= ence, i.e. optional. Hence, [+-]? matches an optional leading = + or - sign.
  3. We have covered three occurrence indicators: + for one or = more, * for zero or more, and ? for zero or one.<= /li>

1.7  Example: Identifiers (or Names) [a-zA-Z_][0-9a-zA-Z_]* or [a-zA-Z_]\w*

  1. Begin with one letters or underscore, followed by zero or more digits, = letters and underscore.
  2. You can use metacharacter \w for a word character= [a-zA-Z0-9_]. Recall that metacharacter \d can be used for a digit [0-9].

1.8  Example: Image Filenames ^\w+\.(gif|png|jpg|jpeg)$

  1. The position anchors ^ and $ match = the beginning and the ending of the input string, respectively. That is, th= is regex shall match the entire input string, instead of a part of the inpu= t string (substring).
  2. \w+ matches one or more word characters (same as [a-= zA-Z0-9_]+).
  3. \. matches the dot (.) character. We need to = use \. to represent . as . has speci= al meaning in regex. The \ is known as the escape code, which = restore the original literal meaning of the following character. Similarly,= *, +, ? (occurrence indicators), ^, $ (position anchors) have special meaning in reg= ex. You need to use an escape code to match with these characters.
  4. (gif|png|jpg|jpeg) matches either "gif", "png", "jpg" or "jpeg". The | denotes "OR" operator. The parentheses are used for grouping the selecti= ons.
  5. The modifier i after the regex specifies case-ins= ensitive matching (applicable to some languages like Perl and JavaScript on= ly). That is, it accepts "test.GIF" and "TesT.Gif= ".

1.9  Example: Email Addresses ^\w+([.-]?\w+)*@\w+([.-]?\w+)*(\.\w{2,3})+$

  1. The position anchors ^ and $ match t= he beginning and the ending of the input string, respectively. That is, thi= s regex shall match the entire input string, instead of a part of the input= string (substring).
  2. \w+ matches 1 or more word characters (same as [a-zA= -Z0-9_]+).
  3. [.-]? matches an optional character . or -. Although dot (.) has special meaning in regex, in = a character class (square brackets) any characters except ^, <= code>-, ] or \ is a literal, and do not re= quire escape sequence.
  4. ([.-]?\w+)* matches 0 or more occurrences of [.-]?\w= +.
  5. The sub-expression \w+([.-]?\w+)* is used to match the use= rname in the email, before the @ sign. It begins with at least= one word character [a-zA-Z0-9_], followed by more word charac= ters or . or -. However, a . or - must follow by a word character [a-zA-Z0-9_]. That = is, the input string cannot begin with . or -; an= d cannot contain "..", "--", ".-" or= "-.". Example of valid string are "a.1-2-3".
  6. The @ matches itself. In regex, all characters other than = those having special meanings matches itself, e.g., a matches = a, b matches b, and etc.
  7. Again, the sub-expression \w+([.-]?\w+)* is used to match = the email domain name, with the same pattern as the username described abov= e.
  8. The sub-expression \.\w{2,3} matches a . foll= owed by two or three word characters, e.g., ".com", ".ed= u", ".us", ".uk", ".co".
  9. (\.\w{2,3})+ specifies that the above sub-expression could= occur one or more times, e.g., ".com", ".co.uk",= ".edu.sg" etc.

Exercise: Interpret this regex, whic= h provide another representation of email address: ^[\w\-\.\+]+\@[a-z= A-Z0-9\.\-]+\.[a-zA-z0-9]{2,4}$.

1.10  Example: Swapping Words using Parenthe= sized Back-References ^(\S+)\s+(\S+)$ and $2 $1

  1. The ^ and $ match the beginning and ending of= the input string, respectively.
  2. The \s (lowercase s) matches a whitespace (bl= ank, tab \t, and newline \r or \n). = On the other hand, the \S+ (uppercase S) matches = anything that is NOT matched by \s, i.e., non-whitespace. In r= egex, the uppercase metacharacter denotes the inverse of the lower= case counterpart, for example, \w for word character and \W for non-word character; \d for digit and \D or non-digit.
  3. The above regex matches two words (without white spaces) separated by o= ne or more whitespaces.
  4. Parentheses () have two meanings in regex:
    1. to group sub-expressions, e.g., (abc)*
    2. to provide a so-called back-reference for capturing and extrac= ting matches.
  5. The parentheses in (\S+), called parenthesized back-re= ference, is used to extract the matched substring from the input strin= g. In this regex, there are two (\S+), match the first two wor= ds, separated by one or more whitespaces \s+. The two matched = words are extracted from the input string and typically kept in special var= iables $1 and $2 (or \1 and \2= in Python), respectively.
  6. To swap the two words, you can access the special variables, and print = "$2 $1" (via a programming language); or substitute operator "= s/(\S+)\s+(\S+)/$2 $1/" (in Perl).
Code Example in Python

Python keeps the parenthesized back references in \1, \2, .... Also, \0 keeps the entire match.

$ python3
>>> re.findall(r'^(\S+)\s+(\S+)$', 'apple orange')
[('apple', 'orange')]     # A list of tuples =
if the pattern has more than one back references
# Back references are kept in \1, \2, \3, etc=
.
>>> re.sub(r'^(\S+)\s+(\S+)$', r'\2 \1', 'apple orange')   # Prefix r for raw string which ign=
ores escape
'orange apple'
>>> re.sub(r'^(\S+)\s+(\S+)$', '\\2 \\1', 'apple orange')<=
/strong>  # Need to use \\ for \ for regular =
string
'orange apple'
Code Example in Java

Java keeps the parenthesized back references in $1, $= 2, ....

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
import java.util.regex.Pattern;
import java.util.regex.Matcher;

public class TestRegexSwapWords {
   public static void main(String[] args) {
      String inputStr =3D "apple orange";
      String regexStr =3D "^(\\S+)\\s+(\\S+)$";  // Regex pattern to be matched
      String replacementStr =3D "$2 $1";  =
       // Replacement pattern with back refer=
ences

      // Step 1: Allocate a Pattern object to=
 compile a regex
      Pattern pattern =3D Pattern.compile(regexStr);

      // Step 2: Allocate a Matcher object fr=
om the Pattern, and provide the input
      Matcher matcher =3D pattern.matcher(inputStr);

      // Step 3: Perform the matching and pro=
cess the matching result
      String outputStr =3D matcher.replaceFirst(replacementStr); // first match only
      System.out.println(outputStr);   // Out=
put: orange apple
   }
}

1.11  Example: HTTP Addresses ^http:\/\/\S+(\/\S+)*(\/)?$

  1. Begin with http://. Take note that you may need to write <= code>/ as \/ with an escape code in some languages (Jav= aScript, Perl).
  2. Followed by \S+, one or more non-whitespaces, for the doma= in name.
  3. Followed by (\/\S+)*, zero or more "/...", for the sub-dir= ectories.
  4. Followed by (\/)?, an optional (0 or 1) trailing /, for directory request.

1.12  Example: Regex Patterns in AngularJS

The following rather complex regex patterns are used by AngularJS in Jav= aScript syntax:

var ISO_DATE_REGEXP =3D /^\d{4,}-[01]\d-[0-3]\=
dT[0-2]\d:[0-5]\d:[0-5]\d\.\d+(?:[+-][0-2]\d:[0-5]\d|Z)$/;

var URL_REGEXP =3D /^[a-z][a-z\d.+-]*:\/*(?:[^:@]+(?::[^@]+)?@)?(?:[^\s:/?#=
]+|\[[a-f\d:]+])(?::\d+)?(?:\/[^?#]*)?(?:\?[^#]*)?(?:#.*)?$/i;

var EMAIL_REGEXP =3D /^(?=3D.{1,254}$)(?=3D.{1,64}@)[-!#$%&'*+/0-9=3D?A=
-Z^_`a-z{|}~]+(\.[-!#$%&'*+/0-9=3D?A-Z^_`a-z{|}~]+)*@[A-Za-z0-9]([A-Za-=
z0-9-]{0,61}[A-Za-z0-9])?(\.[A-Za-z0-9]([A-Za-z0-9-]{0,61}[A-Za-z0-9])?)*$/=
;
// Match both uppercase and lowercase letters=
, single-quote but not double-quote

var NUMBER_REGEXP =3D /^\s*(-|\+)?(\d+|(\d*(\.\d*)))([eE][+-]?\d+)?\s*$/;

var DATE_REGEXP =3D /^(\d{4,})-(\d{2})-(\d{2})$/;

var DATETIMELOCAL_REGEXP =3D /^(\d{4,})-(\d\d)-(\d\d)T(\d\d):(\d\d)(?::(\d\=
d)(\.\d{1,3})?)?$/;

var WEEK_REGEXP =3D /^(\d{4,})-W(\d\d)$/;

var MONTH_REGEXP =3D /^(\d{4,})-(\d\d)$/;

var TIME_REGEXP =3D /^(\d\d):(\d\d)(?::(\d\d)(\.\d{1,3})?)?$/;

1.13  Example: Sample Regex in Perl
s/^\s+//        =
# Remove leading whitespaces (substitute with empty string)
s/\s+$//        # Remove trailing whitespaces=

s/^\s+.*\s+$//  # Remove leading and trailing=
 whitespaces

2.  Regular Expression (Regex) Syntax

A Regular Expression (or Regex) is a pattern (or filter) that describes a set of strings that matches the pattern. In other = words, a regex accepts a certain set of strings and rejects the rest.

A regex consists of a sequence of characters, metacharacters (such as = ., \d, \D, \s, \= S, \w, \W) and operators (such as += , *, ?, |, ^). T= hey are constructed by combining many smaller sub-expressions.

2.1  Matching a Single Character

The fundamental building blocks of a regex are patterns that match a single character. Most characters, including all letters (= a-z and A-Z) and digits (0-9), match itse= lf. For example, the regex x matches substring "x"; z matches "z"; and 9 matches "9".

Non-alphanumeric characters without special meaning in regex also matche= s itself. For example, =3D matches "=3D"; @= matches "@".

2.2  Regex Special Characters and Escape Sequences

Regex's Special Characters

These characters have special meaning in regex (I will discuss in detail= in the later sections):

  • metacharacter: dot (.)
  • bracket list: [ ]
  • position anchors: ^, $
  • occurrence indicators: +, *, ?, = { }
  • parentheses: ( )
  • or: |
  • escape and metacharacter: backslash (\)
Escape Sequences

The characters listed above have special meanings in regex. To match the= se characters, we need to prepend it with a backslash (\), kn= own as escape sequence.  For examples, \+ matche= s "+"; \[ matches "["; and \.<= /code> matches ".".

Regex also recognizes common escape sequences such as \n fo= r newline, \t for tab, \r for carriage-return, \nnn for a up to 3-digit octal number, \xhh for a t= wo-digit hex code, \uhhhh for a 4-digit Unicode, \uhhhhh= hhh for a 8-digit Unicode.

Code Example in Python
$ python3
>>> import re   # N=
eed module 're' for regular expression
# Try find: re.findall(regexStr, inStr) ->=
 matchedStrList
# r'...' denotes raw strings which ignore escape code, i.e., r'\n' is '\'+'=
n'
>>> re.findall(r'a', 'abcabc')
['a', 'a']
>>> re.findall(r'=3D', 'abc=3Dabc')   # '=3D' is not a special regex character
['=3D']
>>> re.findall(r'\.', 'abc.com')  # '.' is a special regex character, need regex escape sequen=
ce
['.']
>>> re.findall('\\.', 'abc.com')  # You need to write \\ for \ in regular Python string
['.']
Code Example in JavaScript

[TODO]

Code Example in Java

[TODO]

2.3  Matching a Sequence of Characters (String or Text)

Sub-Expressions

A regex is constructed by combining many smaller sub-expressions or atoms. For example, the regex Friday matches th= e string "Friday". The matching, by default, is case-sensitive= , but can be set to case-insensitive via modifier.

2.4  OR (|) Operator

You can provide alternatives using the "OR" operator, denoted b= y a vertical bar '|'. For example, the regex four|for|fl= oor|4 accepts strings "four", "for", "floor" or "4".

2.5  Bracket List (Character Class) [...], [^...], [.-.]

A bracket expression is a list of characters enclosed = by [ ], also called character class. It matches ANY O= NE character in the list. However, if the first character of the list is t= he caret (^), then it matches ANY ONE character NOT in the lis= t. For example, the regex [02468] matches a single digit 0, 2, 4, 6, or 8; the regex [^02468] matches any single character other than= 0, 2, 4, 6, or 8= .

Instead of listing all characters, you could use a range expression<= /em> inside the bracket. A range expression consists of two characters sep= arated by a hyphen (-). It matches any single character that = sorts between the two characters, inclusive. For example, [a-d] is the same as [abcd]. You could include a caret (^) in front of the range to invert the matching. For example,= [^a-d] is equivalent to [^abcd].

Most of the special regex characters lose their meaning inside bracket l= ist, and can be used as they are; except ^, -, ] or \.

  • To include a ], place it first in the list, or use escape = \].
  • To include a ^, place it anywhere but first, or use escape= \^.
  • To include a - place it last, or use escape \-.
  • To include a \, use escape \\.
  • No escape needed for the other characters such as ., +, *, ?, (, ), = {, }, and etc, inside the bracket list
  • You can also include metacharacters (to be explained in the next sectio= n), such as \w, \W, \d, \D, \s, \S inside the bracket list.
Name Character Classes in Bracket List (For Perl Only?)

Named (POSIX) classes of characters are pre-defined within bracket expre= ssions. They are:

  • [:alnum:], [:alpha:], [:digit:]:= letters+digits, letters, digits.
  • [:xdigit:]: hexadecimal digits.
  • [:lower:], [:upper:]: lowercase/uppercase let= ters.
  • [:cntrl:]: Control characters
  • [:graph:]: printable characters, except space.
  • [:print:]: printable characters, include space.
  • [:punct:]: printable characters, excluding letters and dig= its.
  • [:space:]: whitespace
=20 =20

For example, [[:alnum:]] means [0-9A-Za-z]. (= Note that the square brackets in these class names are part of the symbolic= names, and must be included in addition to the square brackets delimiting = the bracket list.)

2.6  Metacharacters ., \w, \W, \d, \D, \s, \S

A metacharacter is a symbol with a special meaning inside a reg= ex.

  • The metacharacter dot (.) matches any single character exc= ept newline \n (same as [^\n]). For example, ... matches any 3 characters (including alphabets, numbers, white= spaces, but except newline); the.. matches "there= ", "these", "the  ", and so on.
  • \w (word character) matches any single letter, number or u= nderscore (same as [a-zA-Z0-9_]). The uppercase counterpart <= code>\W (non-word-character) matches any single character that doesn= 't match by \w (same as [^a-zA-Z0-9_]).
  • In regex, the uppercase metacharacter is always the inverse of= the lowercase counterpart.
  • \d (digit) matches any single digit (same as [0-9]). The uppercase counterpart \D (non-digit) matches any s= ingle character that is not a digit (same as [^0-9]).
  • \s (space) matches any single whitespace (same as [ = \t\n\r\f], blank, tab, newline, carriage-return and form-feed). The = uppercase counterpart \S (non-space) matches any single charac= ter that doesn't match by \s (same as [^ \t\n\r\f]).

Examples:

\s\s      # Matc=
hes two spaces
\S\S\s    # Two non-spaces followed by a spac=
e
\s+       # One or more spaces
\S+\s\S+  # Two words (non-spaces) separated =
by a space

2.7  Backslash (\) and Regex= Escape Sequences

Regex uses backslash (\) for two purposes:

  1. for metacharacters such as \d (digit), \D (non-digit), \s (space), \S (non-space), \w (word), \W (non-word).
  2. to escape special regex characters, e.g., \. for ., \+ for +, \* for *, \? for ?. You also need to write \\ for \ in regex to avoid ambiguity.
  3. Regex also recognizes \n for newline, \t for = tab, etc.

Take note that in many programming languages (C, Java, Python), backslas= h (\) is also used for escape sequences in string, e.g., "\n" for newline, "\t" for tab, and you also need to w= rite "\\" for \. Consequently, to write regex pat= tern \\ (which matches one \) in these languages,= you need to write "\\\\" (two levels of escape!!!). Similarly= , you need to write "\\d" for regex metacharacter \d. This is cumbersome and error-prone!!!

2.8  Occurrence Indicators (Repetition Operators): +, *, ?, {m}, {m,n}, {m,}

A regex sub-expression may be followed by an occurrence indicator (aka repetition operator):

  • ?: The preceding item is optional and matched at most once= (i.e., occurs 0 or 1 times or optional).
  • *: The preceding item will be matched zero or more times, = i.e., 0+
  • +: The preceding item will be matched one or more times, i= .e., 1+
  • {m}: The preceding item is matched exactly m times.
  • {m,}: The preceding item is matched m or more times, i.e.,= m+
  • {m,n}: The preceding item is matched at least m times, but= not more than n times.

For example: The regex xy{2,4} accepts "xyy", = "xyyy" and "xyyyy".

2.9  Modifiers

You can apply modifiers to a regex to tailor its behavior, such as globa= l, case-insensitive, multiline, etc. The ways to apply modifiers differ amo= ng languages.

In Perl, you can attach modifiers after a regex, in the form of= /.../modifiers. For examples:

m/abc/i     # ca=
se-insensitive matching
m/abc/g     # global (Match ALL instead of ma=
tch first)

In Java, you apply modifiers when compiling the regex Pattern. For example,

Pattern p1 =3D Pattern.compile(regex, Pattern.=
CASE_INSENSITIVE);  // for case-insensitive m=
atching
Pattern p2 =3D Pattern.compile(regex, Pattern.MULTILINE);         // for multiline input string
Pattern p3 =3D Pattern.compile(regex, Pattern.DOTALL);            // Dot (.) matches all characters including newline

The commonly-used modifer modes are:

  • Case-Insensitive mode (or i): case-insensitive matching fo= r letters.
  • Global (or g): match All instead of first match.
  • Multiline mode (or m): affect ^, $, \A and \Z. In multiline mode, ^ = matches start-of-line or start-of-input; $ matches end-of-line= or end-of-input, \A matches start-of-input; \Z m= atches end-of-input.
  • Single-line mode (or s): Dot (.) will match a= ll characters, including newline.
  • Comment mode (or x): allow and ignore embedded comment sta= rting with # till end-of-line (EOL).
  • more...

2.10  Greediness, Laziness and Backtracking for Repetition Op= erators

Greediness of Repetition Operators *, +, ?, {m,n}: The= repetition operators are greedy operators, and by default grasp a= s many characters as possible for a match. For example, the regex xy= {2,4} try to match for "xyyyy", then "xyyy= ", and then "xyy".

Lazy Quantifiers = *?, +?, ?= ?, {m,n}?, {m,}?, : You can put an extra ? after the rep= etition operators to curb its greediness (i.e., stop at the shortest match)= . For example,

input =3D "The <code>first</code> =
and <code>second</code> instances"
regex =3D <code>.*</code> matches "<code>first</code&g=
t; and <code>second</code>"
But
regex =3D <code>.*?</code> produces two matches: "<code>f=
irst</code>" and "<code>second</code>"

Backtracking: If a regex reaches a s= tate where a match cannot be completed, it backtracks by unwinding one char= acter from the greedy match. For example, if the regex z*zzz i= s matched against the string "zzzz", the z* first= matches "zzzz"; unwinds to match "zzz"; unwinds = to match "zz"; and finally unwinds to match "z", = such that the rest of the patterns can find a match.

Possessive Quantifiers *+, ++, ?+, {m,n}+, {m,}+: You can put an extra + to the rep= etition operators to disable backtracking, even it may result in match fail= ure. e.g, z++z will not match "zzzz". This featur= e might not be supported in some languages.

2.11  Position Anchors ^, $, \b, \B, \<, \>, \A, = \Z

Positional anchors DO NOT match actual character, but matches <= em>position in a string, such as start-of-line, end-of-line, start-of-= word, and end-of-word.

  • ^ and $: The ^ matches the start-of-line. The $ matche= s the end-of-line excluding newline, or end-of-input (for input not ending = with newline). These are the most commonly-used position anchors. For examp= les,
    ing$           # ending with 'ing'
    ^testing 123$  # Matches only one pattern. Sh=
    ould use equality comparison instead.
    ^[0-9]+$       # Numeric string
  • \b and \B: The \b matches the boundary of a word (i.e., start-of-w= ord or end-of-word); and \B matches inverse of \b= , or non-word-boundary. For examples,
    \bcat\b        #=
     matches the word "cat" in input string "This is a cat."
                   # but does not match input "This is a catalog."
  • \< and \>: The \< and <= code>\> match the start-of-word and end-of-word, respectively (co= mpared with \b, which can match both the start and end of a wo= rd).
  • \A and \Z: The \A matches the start of the input. The \Z matches the end of the input.
    They are different from ^ and $ when it comes t= o matching input with multiple lines. ^ matches at the start o= f the string and after each line break, while \A only matches = at the start of the string. $ matches at the end of the string= and before each line break, while \Z only matches at the end = of the string. For examples,
    $ python3
    # Using ^ and $ in multiline mode
    >>> p1 =3D re.compile(r'^.+$', re.MULTILINE)  # . for any character except newline
    >>> p1.findall('testing\ntesting')
    ['testing', 'testing']
    >>> p1.findall('testing\ntesting\n')
    ['testing', 'testing']
       # ^ matches start-of-input or after each l=
    ine break at start-of-line
       # $ matches end-of-input or before line break at end-of-line
       # newlines are NOT included in the matches
    
    # Using \A and \Z in multiline mode
    >>> p2 =3D re.compile(r'\A.+\Z', re.MULTILINE)
    >>> p2.findall('testing\ntesting')
    []    # This pattern does not match the inter=
    nal \n
    >>> p3 =3D re.compile(r'\A.+\n.+\Z', re.MULTILINE)  # to match the internal \n
    >>> p3.findall('testing\ntesting')
    ['testing\ntesting']
    >>> p3.findall('testing\ntesting\n')
    []    # This pattern does not match the trail=
    ing \n
       # \A matches start-of-input and \Z matches=
     end-of-input

2.12  Capturing Matches via Parenthesized Back-References &am= p; Matched Variables $1, $2, ...

Parentheses ( ) serve two purposes in regex:

  1. Firstly, parentheses ( ) can be used to group sub-expressi= ons for overriding the precedence or applying a repetition operator. For e= xample, (abc)+ (accepts abc, a= bcabc, abcabcabc, ...) is different from abc+ (accepts abc, abcc, abccc, ...).=
  2. Secondly, parentheses are used to provide the so called back-refere= nces (or capturing groups). A back-reference contains the matched substring. For examples, the regex (\S+) creat= es one back-reference (\S+), which contains the first word (co= nsecutive non-spaces) of the input string; the regex (\S+)\s+(\S+) creates two back-references: (\S+) and another (\S+= ), containing the first two words, separated by one or more spaces <= code>\s+.

These back-references (or capturing groups) are stored in special varia= bles $1, $2, =E2=80=A6 (or \1, \2, ... in Python), where $1 contains the substring ma= tched the first pair of parentheses, and so on. For example, (\S+)\s+= (\S+) creates two back-references which matched with the first two w= ords. The matched words are stored in $1 and $2 (= or \1 and \2), respectively.

Back-references are important to manipulate the string. Back-references= can be used in the substitution string as well as the pattern. For example= s,

# Swap the first=
 and second words separated by one space
s/(\S+) (\S+)/$2 $1/;                     # P=
erl
re.sub(r'(\S+)  (\S+)', r'\2 \1', inStr)  # P=
ython

# Remove duplicate word
s/(\w+)  $1/$1/;                    # Perl
re.sub(r'(\w+)  \1', r'\1', inStr)  # Python<=
/span>

2.13  (Advanced) Lookahead/Lookbehind, Groupings and Conditio= nal

These feature might not be supported in some languages.

Positive Lookahead (?=3Dpattern)

The (?=3Dpattern) is known as positive lookahead. = It performs the match, but does not capture the match, returning only the r= esult: match or no match. It is also called assertion as it does n= ot consume any characters in matching. For example, the following complex r= egex is used to match email addresses by AngularJS:

^(?=3D.{1,254}$)(?=3D.{1,64}@)[-!#$%&'*+/0=
-9=3D?A-Z^_`a-z{|}~]+(\.[-!#$%&'*+/0-9=3D?A-Z^_`a-z{|}~]+)*@[A-Za-z0-9]=
([A-Za-z0-9-]{0,61}[A-Za-z0-9])?(\.[A-Za-z0-9]([A-Za-z0-9-]{0,61}[A-Za-z0-9=
])?)*$
=20

The first positive lookahead patterns ^(?=3D.{1,254}$) sets= the maximum length to 254 characters. The second positive lookahead = ^(?=3D.{1,64}@) sets maximum of 64 characters before the '@'<= /code> sign for the username.

Negative Lookahead (?!pattern)

Inverse of (?=3Dpattern). Match if pattern is = missing. For example, a(?=3Db) matches 'a' in 'abc' (not consuming 'b'); but not 'acc'. Whereas a(?!b) matches 'a' in 'acc', but not abc.

Positive Lookbehind (?<=3Dpattern= )

[TODO]

Negative Lookbehind (?<!pattern)<= /span>

[TODO]

Non-Capturing Group (?:pattern)

Recall that you can use Parenthesized Back-References to capture the mat= ches. To disable capturing, use ?: inside the parentheses in t= he form of (?:pattern). In other words, ?: disabl= es the creation of a capturing group, so as not to create an unnecessary ca= pturing group.

Example: [TODO]

Named Capturing Group (?<name>= pattern)

The capture group can be referenced later by name.=

Atomic Grouping (>pattern)=

Disable backtracking, even if this may lead to match failure.

Conditional (?(Cond)then|else)

[TODO]

2.14  Unicode

The metacharacters \w, \W, (word and non-word = character), \b, \B (word and non-word boundary) r= econgize Unicode characters.

[TODO]

3.  Regex in Programming Languages

Python: See "Pyt= hon re module for Regular Expression"

Java: See "Regular Expression= s in Java"

JavaScript: See "Regular Expression in JavaScript"

Perl: See "R= egular Expressions in Perl"

PHP: [Link]

C/C++: [Link]

REFERENCES & RESOURCES

  1. (Python) Python's Regular Expression HOWTO @ https://docs.python.org/3/howto/regex.html= (Python 3).
  2. (Python) Python's re - Regular expression operations @ https://docs.python.org/3/library/re.= html (Python 3).
  3. (Java) Online Java Tutorial's Trail on "Regular Expressions" @ htt= ps://docs.oracle.com/javase/tutorial/essential/regex/index.html.
  4. (Java) JavaDoc for java.util.regex Package @ https://docs.oracle.com/javase/10/docs/api/java/util/regex/package-summ= ary.html (JDK 10).
  5. (Perl) perlrequick - Perl regular expressions quick start @ https://perldoc.perl.org/perlreq= uick.html.
  6. (Perl) perlre - Perl regular expressions @ https://perldoc.perl.org/perlre.html.
  7. (JavaScript) Regular Expressions @ https://developer= .mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_Expressions.

Last modified: November, 2018

Feedback, comments, correctio= ns, and errata can be sent to Chua Hock-Chuan (ehchua@ntu.edu.sg)  &nb= sp;|   HOME

------MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD---- Content-Type: text/css Content-Transfer-Encoding: quoted-printable Content-Location: https://www3.ntu.edu.sg/home/ehchua/programming/css/programming_notes_v1.css @charset "utf-8"; * { margin: 0px; padding: 0px; } body { background-color: rgb(202, 251, 237); color: rgb(0, 0, 0); font-fami= ly: "Segoe UI", Segoe, Calibri, "Nimbus Sans L", Ubuntu, Tahoma, Arial, Hel= vetica, Verdana, sans-serif; font-size: 14px; text-align: justify; line-hei= ght: 1.5; } #wrap-outer { margin: 20px; padding: 0px; } #wrap-inner { background-color: rgb(255, 255, 255); margin: 0px; border: 1p= x solid rgb(221, 221, 221); padding: 25px 15px; box-shadow: rgb(221, 221, 2= 21) 5px 5px 0px; } #content-header { margin: 0px; padding: 50px 0px 10px; } #content-main { margin: 0px; padding: 30px 0px 20px; } #content-footer { font-family: "Century Gothic", "Segoe UI", Segoe, "Nimbus= Sans L", Ubuntu, Verdana, Tahoma, Arial, Helvetica, sans-serif; font-size:= 13px; text-align: right; color: rgb(192, 80, 77); margin: 30px 0px 0px; pa= dding: 0px; border-top: 4px solid rgb(12, 155, 116); } .header-footer { font-family: "Century Gothic", "Segoe UI", Segoe, "Nimbus = Sans L", Ubuntu, Verdana, Tahoma, Arial, Helvetica, sans-serif; color: rgb(= 192, 80, 77); font-size: 13px; text-align: right; margin: 10px 0px 5px; pad= ding: 5px 4px; } .header-footer a { color: rgb(192, 80, 77); text-decoration: underline; } .header-footer a:focus, .header-footer a:hover { text-decoration: none; col= or: rgb(11, 83, 149); } h1, h2, h3, h4, h5, h6 { font-family: "Century Gothic", "Trebuchet MS", "Se= goe UI", Segoe, "Nimbus Sans L", Ubuntu, Verdana, Tahoma, Arial, Helvetica,= sans-serif; margin: 0px; color: rgb(10, 132, 100); letter-spacing: 1px; li= ne-height: 1.2; text-align: left; } h1 { font-size: 40px; font-weight: 400; padding: 0.2em 0px; } h2 { font-size: 36px; font-weight: 400; padding: 0.2em 0px; } h3 { font-size: 22px; border-bottom: thin solid rgb(12, 155, 116); padding:= 1.5em 0px 0.3em; } h4 { font-family: "Segoe UI", Segoe, "Nimbus Sans L", Ubuntu, Verdana, Taho= ma, Arial, Helvetica, sans-serif; font-size: 18px; padding: 1.3em 0px 0.2em= ; border-bottom: thin dotted rgb(12, 155, 116); } h5, h6 { font-family: "Segoe UI", Segoe, "Nimbus Sans L", Ubuntu, Verdana, = Tahoma, Arial, Helvetica, sans-serif; font-size: 15px; color: rgb(68, 68, 6= 8); padding: 1.2em 0px 0px; letter-spacing: 1px; } .line-heading { color: rgb(68, 68, 68); font-size: 15px; font-weight: bold;= letter-spacing: 1px; padding: 0.2em 0px; } .line-heading-code-new { font-family: Consolas, "DejaVu Sans Mono", "Lucida= Console", "Courier New", Courier, monospace; color: rgb(227, 27, 35); font= -size: 15px; font-weight: normal; } p { margin-top: 0.6em; margin-bottom: 0.4em; } pre { font-family: Consolas, "DejaVu Sans Mono", "Lucida Console", "Courier= New", Courier, monospace; font-size: 13px; margin: 5px 0px 8px; border: 2p= x solid rgb(248, 248, 248); padding: 5px 10px; line-height: 135%; } code { font-family: Consolas, "DejaVu Sans Mono", "Lucida Console", "Courie= r New", Courier, monospace; } ul { margin: 0.3em 0px 0.2em 1.8em; padding: 0px; list-style-image: url("im= ages/BulletSquare.png"); } ul ul li { list-style-image: url("images/BulletRound.png"); } ul ul u1 li { list-style-type: circle; list-style-image: none; } ol { list-style-type: decimal; margin: 0.3em 0px 0.2em 2.5em; padding: 0px;= } ol ol li { list-style-type: lower-alpha; } ol ol o1 li { list-style-type: lower-roman; } li { margin: 0.4em 0px; } .float-left-ol-ul { overflow: hidden; } .float-left-li { position: relative; left: 20px; margin-right: 20px; } a { color: rgb(11, 83, 149); text-decoration: none; } a:hover, a:focus { color: rgb(192, 80, 77); text-decoration: underline; } a.references { display: block; width: 30em; font-size: 18px; font-weight: b= old; margin: 4em 0px 0px; } p.references { font-size: 18px; font-weight: bold; margin: 4em 0px 0px; } .center-block { margin: 10px auto; } .text-center { text-align: center; } .text-right { text-align: right; } .underline { text-decoration: underline; } .font-code { font-family: Consolas, "DejaVu Sans Mono", "Lucida Console", "= Courier New", Courier, monospace; } .font-code-text { font-family: Consolas, "DejaVu Sans Mono", "Lucida Consol= e", "Courier New", Courier, monospace; font-size: 14px; } .font-code-smaller { font-family: Consolas, "DejaVu Sans Mono", "Lucida Con= sole", "Courier New", Courier, monospace; font-size: 13px; } .font-normal { font-family: "Segoe UI", Segoe, Calibri, "Nimbus Sans L", Ub= untu, Tahoma, Arial, Helvetica, Verdana, sans-serif; } .pre { white-space: pre; } .color-example { background-color: rgb(215, 236, 211); } .color-example-light { background-color: rgb(236, 246, 234); } .color-syntax, .color-command { background-color: rgb(204, 238, 241); } .color-explanation { background-color: rgb(238, 238, 238); } .color-comment { color: rgb(0, 153, 0); } .color-output { color: rgb(0, 82, 162); } .color-new { color: rgb(227, 27, 35); } .color-error { color: rgb(231, 84, 128); } .color-plain { background-color: rgb(255, 255, 255); } .color-highlight { background-color: rgb(238, 238, 68); } .color-highlight-new { background-color: rgb(255, 255, 204); } .output { background-color: rgb(236, 246, 234); border: 2px solid rgb(248, = 248, 248); padding: 4px 8px; } .side-note { margin-top: 15px; margin-left: 40px; padding: 3px 8px; backgro= und-color: rgb(231, 231, 231); } img.image-center { display: block; margin: 10px auto; } img.image-border { border: thin solid rgb(221, 221, 221); } img.image-float-left { float: left; margin: 8px 15px 15px 0px; border: thin= solid rgb(221, 221, 221); } img.image-float-right { float: right; margin: 8px 0px 15px 15px; border: th= in solid rgb(221, 221, 221); } .float-clear { clear: both; } .table-zebra, .table-program { border-collapse: collapse; border: 0px; marg= in: 0px auto; padding: 0px; width: 100%; background-color: rgb(231, 240, 24= 8); text-align: left; vertical-align: top; } .table-zebra tr > th { color: rgb(255, 255, 255); background-color: rgb(0, = 157, 217); margin: 0px; border: 2px solid white; padding: 4px 10px; font-si= ze: 15px; letter-spacing: 1px; text-align: center; } .table-zebra tr > td { margin: 0px; border: 2px solid white; padding: 2px 8= px; vertical-align: top; } .table-zebra tr:nth-child(2n+1) > td { background-color: rgb(203, 223, 241)= ; } td > pre { font-family: Consolas, "DejaVu Sans Mono", "Lucida Console", "Co= urier New", Courier, monospace; font-size: 14px; margin: 0px; border: none;= padding: 2px 0px 5px; line-height: 135%; } .table-program th { color: rgb(255, 255, 255); background-color: rgb(0, 157= , 217); margin: 0px; border: 2px solid white; padding: 4px 10px; font-size:= 15px; letter-spacing: 1px; text-align: center; } .table-program td { margin: 0px; border: 0px; padding: 0px; } .table-program td pre { margin: 0px; border: 2px solid rgb(248, 248, 248); = padding: 5px 10px; } .table-program td pre.text-right { text-align: right; } .col-desc { background-color: rgb(231, 240, 248); } .col-code, .tr-alt { background-color: rgb(203, 223, 241); } .col-example { background-color: rgb(238, 238, 238); } .col-line-number { width: 40px; background-color: rgb(225, 233, 207); } .col-program { background-color: rgb(240, 244, 233); } #wrap-toc { display: block; background: none 0px 0px repeat scroll rgb(231,= 246, 239); float: right; width: 230px; z-index: 100; line-height: 1.5; mar= gin: 0px 0px 0px 15px; padding: 5px 8px 10px; text-align: left; white-space= : nowrap; } #wrap-toc h5 { letter-spacing: 1px; margin: 0px; text-transform: uppercase;= color: rgb(68, 68, 68); padding: 0.5em 0px; } a#show-toc { color: rgb(192, 80, 77); text-decoration: none; letter-spacing= : 1px; } #toc { overflow: auto; } #toc a.toc-H3 { margin-left: 0px; font-size: 15px; } #toc a.toc-H4 { margin-left: 20px; font-size: 14px; } #toc a.toc-H5 { margin-left: 40px; font-size: 14px; } ------MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD---- Content-Type: image/png Content-Transfer-Encoding: base64 Content-Location: https://www3.ntu.edu.sg/home/ehchua/programming/css/images/BulletSquare.png iVBORw0KGgoAAAANSUhEUgAAAAkAAAANAQMAAABBztZFAAAABlBMVEUAAwBjjJzG2b5OAAAAAXRS TlMAQObYZgAAABBJREFUCNdjYMAG7FARAwMADXkBNzRuJgIAAAAASUVORK5CYII= ------MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD---- Content-Type: image/png Content-Transfer-Encoding: base64 Content-Location: https://www3.ntu.edu.sg/home/ehchua/programming/css/images/BulletRound.png iVBORw0KGgoAAAANSUhEUgAAAAUAAAANAQMAAABb8jbLAAAABlBMVEX///8AUow5QSOjAAAAAXRS TlMAQObYZgAAABNJREFUCB1jYEABBQw/wLCAgQEAGpIDyT0IVcsAAAAASUVORK5CYII= ------MultipartBoundary--fHajKsaKJqK90oMbFBflWc7w3WBKsRfNCAlJ0LUAcD------