This repository contains three compiler course projects that build progressively from basic parsing exercises to a MiniJava compiler that generates LLVM IR.
The codebase is organized as follows:
Project 1 - Compilers: introductory parsing and translation exercises.Project 2 - Compilers: MiniJava parsing, symbol table construction, semantic analysis, and offset computation.Project 3 - Compilers: MiniJava-to-LLVM code generation built on top of the frontend from Project 2.
Taken together, the projects show the usual compiler pipeline in increasing complexity: parsing, semantic analysis, layout computation, and code generation.
Project 1 contains two introductory language-processing exercises.
This part implements a handwritten recursive-descent parser and evaluator for arithmetic expressions over single-digit integers, parentheses, and the operators +, -, *, and /.
The evaluator reads from standard input one character at a time and uses a single lookahead token. Parsing and evaluation happen together: instead of constructing an abstract syntax tree, each grammar rule computes and returns its result directly. The implementation separates expression parsing by precedence, with distinct routines for expressions, terms, and factors.
A notable limitation is that numbers are restricted to a single digit, since digits are processed character by character.
This part defines a small string-processing language and translates programs written in it into Java source code.
The lexical analyzer is implemented in JFlex, a lexer generator that produces tokenizers from lexical rules. The parser is implemented in Java CUP, a parser generator for Java that builds parsers from grammar specifications. In this project, the parser performs syntax-directed translation directly in the grammar actions, producing Java source fragments as parsing proceeds.
The generated output is a Java program containing a Main class, a Java main method, and static helper methods corresponding to functions in the source language. This part demonstrates a complete generated frontend pipeline: lexical analysis with JFlex, parsing with CUP, and immediate source-to-source translation into Java.
Project 2 moves to MiniJava, a small Java-like teaching language, and implements the frontend of a compiler.
The parser is defined in JavaCC, a parser generator that produces a Java parser from a grammar specification. JTB is used on top of that grammar to generate the abstract syntax tree classes and visitor infrastructure, making it easier to implement later compiler passes cleanly.
This stage of the project performs the core frontend tasks of compilation:
- parsing MiniJava source code
- building a symbol table for classes, methods, fields, and variables
- checking semantic correctness and type consistency
- computing field and method offsets needed for object layout and dynamic dispatch
The implementation is organized as multiple passes over the syntax tree. One visitor builds and validates the symbol table, while another performs type checking over statements and expressions. The symbol table also tracks inheritance relationships and is used to compute the memory layout of class fields and method slots.
Project 3 extends the MiniJava frontend into a backend that generates LLVM IR.
LLVM is a compiler infrastructure framework, and LLVM IR is its intermediate representation: a low-level, typed program form that sits between the source language and final machine code. Instead of producing assembly directly, this project translates MiniJava programs into LLVM IR.
This stage reuses the parsing and symbol-table infrastructure from Project 2, but adds the information needed for code generation, such as stored field offsets, method offsets, and class layout data.
Code generation is implemented as a visitor over the syntax tree. It emits LLVM instructions for expressions, statements, object allocation, array operations, control flow, and virtual method calls. The generated code also includes class vtables and runtime helper routines, allowing MiniJava objects and dynamically dispatched methods to be represented explicitly at the IR level.
The projects are intended to be built and run from the command line with the tools used in the course assignments.
- Project 1 Part 1 is a plain Java program and can be compiled and run directly.
- Project 1 Part 2 uses JFlex, a lexer generator, and Java CUP, a parser generator, to build the translator.
- Projects 2 and 3 use JavaCC, a parser generator, together with JTB, which generates syntax-tree classes and visitor scaffolding from the grammar.
The makefile files in the MiniJava projects show the expected build pipeline and assume that the required .jar files are available in the surrounding project directories.
For Projects 2 and 3, input programs are expected under a sibling input/ directory, since both Main.java files prepend ../input/ to the filenames passed on the command line. Project 3 writes generated LLVM files under ../output/.