GATE 2026 CS (CS2) – Question 35
A lexical analyzer uses the following token definitions:
- $letter \to [A-Za-z]$
- $digit \to [0-9]$
- $id \to letter(letter \mid digit)^{*}$
- $number \to digit^{+}$
- $ws \to (blank \mid tab \mid newline)^{+}$
For the string `x1 23mm 78 y 7z zz5 14A 8H AaYcD`, the number of tokens (excluding $ws$) produced by the lexical analyzer is __________. (answer in integer)
Practise this question in The GATE Grind →
Show answer and explanation
Correct answer: 13
Explanation
Using longest-match tokenization, the string is split as follows: `x1` | `23` `mm` | `78` | `y` | `7` `z` | `zz5` | `14` `A` | `8` `H` | `AaYcD`. Counting these non-whitespace tokens gives 13 tokens in total. Hence the answer is 13.