1.Introduction🔗
Tetral, from the Greek word for "four," is a scripting language intended to be interpreted to provide a platform-agnostic execution system and which is intended to be executed fast with little checks on the IR past the initial compilation.
1.1.Core Principles🔗
Additionally, Tetral is designed with the following 4 core principles in mind, hence the name.
- Statically Typed
- Types must be specified and known at compile-time to allow for the compiler to optimize as much as possible and to provide easy readability for programmers.
- Compiled
- Compiled to an intermediate representation (also referred to as a bytecode,) which can be executed on any platform as long as an interpreter exists for the platform.
- Interpreted
- Interpreted by a native interpreter which performs next to checks on the bytecode and executes it to ensure fast execution speed.
- Platform-Agnostic
- As long as an interpreter exists, it shouldn't matter how the IR was compiled, it should be possible to execute it.
1.2.Other Principles🔗
There are more principles this language was designed around, but they aren't as flashy.
- No Memory Interaction
- A programmer should never have to care about memory or be able to interact with raw memory.
- Explicit Null Control
This idea is that object types, should never be null, unless a null value is explicitly allowed. For example, if a variable is declared of an object type, it should be initialized to a value, rather than null. And you shouldn't be able to assign a null value to a non-null type.
The exception for this are arrays and strings, which are null initialized. But null arrays and strings are treated as just a regular array that has a length of 0.
This only goes for object types, primitives are initialized to a zero value, while strings and arrays always null-initialized by default, but null arrays and strings are just treated as having a length of
0.- Reference Counted
- No garbage collector, just reference counting of objects. When the
reference count of an object reaches
0, it is freed and all its properties, if object types as well, have their reference counters decremented.
2.Scope🔗
This document defines the Tetral (Also referred to as TetralScript, or TetralLang) programming language and the Tetral Bytecode format.
3.Concepts🔗
- Source File
- A UTF-8 encoded text file that contains Tetral source code.
- Bytecode/IR/Intermediate Representation
- These terms will be used interchangeably in this document, but they all refer to the compiled representation of Tetral source code.
- Program
- Either a valid Tetral source file or a Tetral bytecode file that can be executed by the interpreter.
- Symbol
A symbol is an object which can be referred to with a specified name. Examples of symbols include, but are not limited to:
- Functions
- Structs
- Enums
- Classes
- Local Variables
- Global Variables
- Imported Symbols
- Namespace
- A namespace is a list of names delimited by :: characters.
- Simple Name
- A simple name is the name of a symbol in a source file with no namespace prefix.
- Fully Qualified Name
- A fully qualified name is the name of a symbol with a namespace prefixed to it.
4.Data Types🔗
A data type is a specific unit of memory that holds data in a specific layout to represent some piece of information.
4.1.Type Categories🔗
Values types are types whose value is stored on the stack and whose lifetime depends on the current stack scope.
Reference types are types which exist in the heap and the user only receives a reference to that type's data on the heap.
4.2.The Boolean Type🔗
The boolean type represents a false/true state.
4.3.Integral Types🔗
Integral types are arithmetic value types of varying sizes used to represent whole numbers. Integral types may be signed or unsigned.
Tetral provides the following integral types:
| Type Name | Size and alignment | Description | Minimum value | Maximum value |
|---|---|---|---|---|
uint8 | 1 | Unsigned 8-bit integer | 0 | 255 |
int8 | 1 | Signed 8-bit integer | -128 | 127 |
uint16 | 2 | Unsigned 16-bit integer | 0 | 65535 |
int16 | 2 | Signed 16-bit integer | -32768 | 32767 |
uint32 | 4 | Unsigned 32-bit integer | 0 | 4294967295 |
int32 | 4 | Signed 32-bit integer | -2147483648 | 2147483647 |
uint64 | 8 | Unsigned 64-bit integer | 0 | 18446744073709551615 |
int64 | 8 | Signed 64-bit integer | -9223372036854775808 | 9223372036854775807 |
4.4.Floating Point Types🔗
The floating-point types are arithmetic types which hold a number value according to the corresponding IEEE-745 standard format.
| Type Name | Size and alignment | Description | IEEE Format |
|---|---|---|---|
float32 | 4 | 32-bit floating point type | binary32 |
float64 | 8 | 64-bit floating point type | binary64 |
4.5.Array Types🔗
Array types are reference types which hold a sequential list of a specific data type.
Each unit of data contained in an array is referred to as an "element," and the amount of items the array holds is called its "length."
An array's length cannot be changed after creation, but all items within an array are mutable unless declared otherwise by a user.
4.6.The String Type🔗
The String type is a reference type which holds a sequential list of UTF-8 code units.
The string type is immutable, none of its elements can be changed. The amount of UTF-8 code units stored in a string is called its "length."
4.7.Struct Types🔗
Struct types are value types made up of "properties," each of which has a name and a pre-set type, determined by the user.
Structs cannot hold any properties that directly reference itself.
4.8.Enum Types🔗
Enum types are value types that declare a list of named constant values with the first always starting at value 0 and each subsequent value having a value one higher than the last.
The name for the underlying value of an enum's named constant is its "ordinal" value.
4.9.Variant Types🔗
Variant types are a value type which define a list of variants the variant type itself can take the form of.
Each variant of a variant type can declare 0 or more properties, each of which has a name and a pre-set type.
The index of a variant in a variant type is referred to as its "ordinal" value.
Variant types can hold type parameters.
4.10.Interface Types🔗
An interface type is a reference type, but not a type which can itself be instantiated.
Interfaces declare 0 or more methods which must be defined by any class wishing to implement the interface.
Implementing an interface means declaring a class as being an implementer of an interface type and defining a function body for each of the methods the interface declares.
Interfaces can hold type parameters.
4.11.Class Types🔗
Class types are reference types which define class members. Class members are either constructors, methods or fields.
Methods are functions that exist within the namespace of a class. Methods are also divided into two separate types: static and non-static. Static methods belong to the class, while non-static methods belong to a class instance.
Class members can, like methods, be declared either static or non-static in which case the same rules for ownership apply.
Members are like the properties of structs and as such they are subject to the same alignment and padding behaviors as defined for struct properties.
Classes can extend one other class and can implement as many non-conflicting interfaces as described in the limits section (See 5. Limits)
Class types can hold type parameters.
4.12.Nullables🔗
A nullable type is a value type which marks another type as nullable, meaning its value may not exist.
4.13.Type Compatibility And Assignability🔗
Type compatibility determines if a type can be passed to a function, set as the value of a variable, a variant/struct property, or a class field. The following table only determines semantic compatibility, it does not list conversions that may or not need to be performed when assigning a value of one type to an object of another, compatible, type.
| Type | Can be assigned from |
|---|---|
| Strings | Only other Strings. |
| Arrays | Only arrays of the same component type. |
| Unsigned Integers | Any unsigned integer of the same or smaller size. |
| Signed Integers | Any signed integer of the same or smaller size, or any unsigned integer of a smaller size. |
| Floating point numbers | Any floating point of the same or smaller size, or any integer. |
| Structs | Only other values of the same struct type. |
| Enums | Only other values of the same enum type. |
| Variants | Only other values of the same variant type. |
| Interfaces | Any class which implements the interface. |
| Classes | Any value of the same class or a derived class. |
| Nullables | Only values of the same nullable type. |
| References | Only values of the same reference type. |
4.14.Type Conversions🔗
This section defines the conversion operations that must occur between types. If there is no conversion listed from one type to another in this section, then it means that nothing must be done to convert from one type to another, or that the conversion is not possible. See the previous section for type compatibility and assignability.
4.14.1.Implicit Conversions🔗
These conversions must be done implicitly without requiring an explicit cast operation by a user of the language. When the term "usual arithmetic conversions" is referenced, the following conversions are what they reference.
| Source Type | Target Type | Conversion |
|---|---|---|
| Any signed or unsigned integral type | A larger signed or integral type | The value is converted from the smaller to the equivalent value in the larger type. |
| Any signed or unsigned integral type | A floating point type | The value is converted from the integral type the closest possible floating point value. |
| A 32-bit floating point value | A 64-bit floating point value | The value is converted to the closest value possible in the larger type. |
4.14.2.Explicit Conversions🔗
These conversions must be explicitly done by users by using a casting expression (Defined in the language grammar).
| Source Type | Target Type | Conversion |
|---|---|---|
| Any integral type | Any integral type of smaller size | The resulting value is the larger integral's value reduced modulo 2N, where N is the number of value bits in the target type. The resulting bit pattern is interpreted according to the target type's signedness. |
| A 64-bit floating point value | A 32-bit floating point value | The larger type's value is converted to the closest equivalent value of the smaller type. |
| Any floating point value | Any integral value | The floating point value is first rounded down, then converted to the equivalent integral value. |
4.14.3.Boolean Conversion🔗
The following table defines how values of different types are to be converted to boolean values. Unless it is stated that a type cannot be converted to a boolean value in the following table, then any value of that type is "boolean-assignable."
| Value | Boolean Result | |
|---|---|---|
| Boolean value | No conversion needed. | |
| Any arithmetic type | 0 values evaluate to false, non-zero values evaluate to true. | |
| strings/arrays | If the array/string has a length of 0, evaluates to false. Otherwise, evaluates to true. | |
| Enums | Enum values with the ordinal value 0 evaluate to false, while non-zero ordinal values evaluate to true. | |
| Variants | Variants with the ordinal 0 evaluate to false, while other ordinals evaluate to true. | |
| Interfaces | Cannot be evaluated as a boolean value. | |
| Classes | ||
| Structs | ||
| Nullables | Evaluate to a boolean value corresponding to the presence state of the nullable. | |
4.15.Type Parameters🔗
Type parameters allow for types to be specified with incomplete types, allowing them to be filled in later.
4.16.Type Comparability🔗
This section defines which types are compatible with each other. For comparison semantics see 6.2. Comparison Semantics. Thus, this section also defines the term "comparable type."
If a data type is not listed here, it cannot be compared to any other type.
- Booleans.
- Arithmetic Types (integral and floating-point types.)
- Arrays which hold a comparable type are comparable.
- Strings.
- Enums.
- Variants.
5.Limits🔗
The following limits are defined for features and aspects of the Tetral Language.
| Aspect | Limit |
|---|---|
| Maximum number of implemented interfaces | 16 |
| Maximum number of function/method/constructor declaration arguments | 62 |
| Maximum number of function/method/constructor call arguments | 62 |
| Maximum number of enum constants per enum type | 1024 |
| Maximum number of variants per variant type | 256 |
| Maximum call stack depth | 512 |
| Maximum number of nested functions | 8 |
| Maximum length of an identifier | 256 |
6.Language🔗
Tetral files are written in UTF-8 encoded text files, referred to as "source files." They must follow the grammar defined in the following sections.
6.1.Identifier Semantics🔗
An Identifier can denote a function, a struct, an interface, an enum, a class, a local variable or a global variable. An identifier with an identical value can reference a different symbol at different points in a program.
A difference must be made between the identifiers used in
Symbol lookup is performed in ascending scopes from the current scopes. The order being: Local Scope, Current Function Scope, Global Scope.
6.2.Comparison Semantics🔗
This section details the semantics for comparing two values of specific types by considering two hypothetical operands: left and right.
6.2.1.Arithmetic Types🔗
Constraints
- The right operand must be assignable to the left operand's type.
Evaluation
- The usual arithmetic conversions are performed.
- If the left operand's value is lesser than the right operand's, then the right operand is greater.
- If the left operand's value is greater than the right operand's, then the left operand is greater.
- If both operands have an equal value, neither is greater.
6.2.2.Array And String Types🔗
Constraints
- Both operands must be arrays of the same, comparable type.
Evaluation
- If the left operand's length is 0, then the right operand is greater.
- If the right operand's length is 0, then the left operand is greater.
- If all elements in both operands are equal, neither operand is greater or lesser, they are equal.
- If the first non-equal value in the left operand is lesser than the equivalent value in the right operand, the right operand is greater.
- If the first non-equal value in the left operand is greater than the equivalent value in the right operand, the left operand is greater.
6.2.3.Struct, Variant and Class Types🔗
Constraints
- Struct, Variant and Class types cannot be compared.
6.2.4.Enum Types🔗
Constraints
- Both operands must be of the same enum type.
Evaluation
- Enum comparison is performed identically to arithmetic type comparison, with the ordinal value of both enums being used to compare them.
6.2.5.Nullable Types🔗
Constraints
- Both operands must hold the same type of value.
- Both operands must hold a comparable value.
Evaluation
- If bother operands have a presence state of false, they are equal and neither is greater.
- If either operand has a false presence state, that operand is lesser.
- The operands are then compared by their held value.
6.3.Equality Semantics🔗
This section details the semantics for checking the equality of two values of specific types by considering two hypothetical operands: left and right.
6.3.1.Arithmetic/Boolean Types🔗
Constraints
- The right operand must be assignable to the left operand's type.
Evaluation
- The usual arithmetic conversions are performed.
- If the values of both operands are equal, they are equal.
6.3.2.Array And String Types🔗
Constraints
- Both operands must be of the same type.
Evaluation
- If the operand's lengths differ, the operands are not equal.
- If all component values in both operands are equal, they are equal. Otherwise, they are not equal.
6.3.3.Struct Types🔗
Constraints
- Both operands must be of the same type.
Evaluation
- The value of each property of both operands are checked for equality.
- If all properties in both operands are equal, the operands are equal. Otherwise, they are not equal.
6.3.4.Class/Interface Types🔗
Constraints
- The right operand must be assignable to the left operand's type.
Evaluation
- If the operands are not of the same type, they are not equal.
- The value of each property of both operands are checked for equality.
- If all properties in both operands are equal, the operands are equal. Otherwise, they are not equal.
6.3.5.Enum Types🔗
Constraints
- Both operands must be of the same enum type.
Evaluation
- For two enum values to be equal, their ordinal values must be equal.
6.3.6.Variant Types🔗
Constraints
- Both operands must be of the same variant type.
Evaluation
- If the operands' variant ordinal values are not the same, they are not equal.
- Each variant property is then checked for equality in both operands.
- If all properties are equal, then the variant values are equal, otherwise they are not.
6.3.7.Nullable Types🔗
Constraints
- Both operands must hold the same type of value.
Evaluation
- If bother operands have a presence state of false, they are equal.
- If either operand has a false presence state, the operands are not equal.
- The operands then have their equality checked by their held value.
6.4.Overflow/Underflow Semantics🔗
Overflows and underflows happen when an arithmetic type's value exceeds its maximum representable value, or when its value subceeds it's minimum representable value.
In such cases, the value wraps around without any errors.
In cases where overflow or underflow is detectable by a compiler, the user should be warned about it.
6.5.Default Initialization🔗
Default initialization occurs when a variable or class/struct property is declared with no value.
Arithmetic types are initialized to the type's equivalent of the value
0. Booleans to the value false. Array, string, nullable
reference types and nullable value-reference-types are initialized to null.
Enumeration types are initialized to the enum value with the ordinal
0.
Variants are initialized to the variant with the ordinal 0. If the variant with value 0 requires arguments, then the type cannot be default-initialized.
Structs are initialized with all properties set to their default values as specified in the struct declaration.
If a type cannot be default initialized, such as in the case of a class with no no-argument constructors, variants where the 0 ordinal variant requires property values to be specified or an interface type. Then the declaration is considered invalid and a value must be explicitly set.
6.6.String Conversion Semantics🔗
This section details how certain types are converted to their string representations. The string data type itself is omitted, as it does not need to be converted to a string.
6.6.1.Booleans🔗
Booleans evaluate to a string with the value "true" or "false" depending on the value of the boolean.
6.6.2.Arithmetic Types🔗
Arithmetic types are converted to string representation of their numerical value in base 10.
For floating point, if required, scientific notation may be used.
6.6.3.Arrays🔗
Arrays evaluate to a comma-delimited string prefixed and suffixed with the "[" and "]" characters respectively.
Empty arrays evaluate to just "[]".
6.6.4.Structs and Classes🔗
Structs and classes evaluate to a string prefixed with the simple name of the type, followed by the "{" character.
Each property is then listed with its name, followed by a ": " sequence. After the delimiter, the property's value is converted to a string and appended. Each property's string representation is delimited with a "," string. If the property is a string, the string value must be prefixed and suffixed with double quotes.
Private fields on classes are included in the string.
6.6.5.Variant Types🔗
Variant types are prefixed with the simple name of the variant type, followed by a "." and the simple name of the variant.
If the variant has declared properties, then append "(" to the string and append a string representation of each property's string representation to the string, each property delimited with a ", " sequence. The property strings are then followed by a ")".
6.6.6.Enum Types🔗
Values of an enum type are converted to string by using the enum constant's name as the string representation.
6.6.7.Nullable Type🔗
If the value is a nullable-reference or a value-type-reference and the pointer is null, then the string representation becomes "null".
If the nullable's presence state is false, then the string representation is "none". Otherwise, the string representation is the present value's string representation.
6.7.Notation🔗
References to other syntactic rules are in italic, literal values are
in
Indicates a syntax element enclosed with "()" which might contain an expression, or several expressions.
6.8.Lexical Elements🔗
Lexical Analysis Rules
- Lexical analysis shall consume the largest sequence of source characters that constitutes a valid token.
- Tokens, outside of character or string literals, ignore all white space.
White space is defined as any Unicode codepoint with the
White_Spacebinary property. - Any use of a literal (eg:
const ) in the definition of the language grammar is case-sensitive and thus must match the keyword exactly. - No Unicode normalization is performed on the parsed tokens.
6.8.1.Keywords🔗
Syntax
Lexical Analysis Rules
- When lexical analysis encounters a sequence which can be both a keyword or an identifier, then keywords take precedence over identifiers.
- After lexical analysis, alias keywords shall have the same syntactic meaning as their corresponding real keywords. For alias token mappings, see Table 8 - Alias token mappings.
- Contextual keywords are only keywords in specific contexts. Those contexts
are defined in the grammar as they appear. Outside the context which will
be defined,
contextual-keywords 's are identifiers.
| Alias Token | Mapped Token |
|---|---|
boolean | bool |
byte | int8 |
ubyte | uint8 |
char | uint8 |
short | int16 |
ushort | uint16 |
int | int32 |
uint | uint32 |
long | int64 |
ulong | uint64 |
6.8.2.Identifiers🔗
Syntax
Constraints
- Identifiers must reference a valid symbol that is accessible in the current semantic context. (Lookup semantics defined in 6.1. Identifier Semantics.)
- An identifier cannot be longer than the defined maximum length of an identifier. (Defined in 5. Limits.)
Semantics
- A symbol lookup is performed as defined in 6.1. Identifier Semantics.
- The result of an identifier expression, is value of the symbol it is referencing.
- An identifier which references a valid semantic variable or a constant is a valid lvalue.
6.8.3.Constants🔗
Syntax
6.8.3.1.Integer Constants🔗
Syntax
Constraints
- Unsuffixed integer constants cannot represent a value that exceeds the maximum value of an unsigned 64-bit integer.
- Suffixed integer constants cannot represent a value that exceeds the suffix specifier's type. (See Table 9 - Suffix type mappings.)
- Any integer constant with the value of
-0 is invalid.
Semantics
- Unsuffixed integer constants result in one of three types of integral types based on the smallest type that can hold the constant value. (See Table 10 - Unsuffixed integer types.)
- Suffixed integer constants result in a value with the integer type specified by the suffix. (See Table 9 - Suffix type mappings.)
| Suffix | Mapped Type |
|---|---|
int8 | |
int16 | |
int32 | |
int64 | |
uint8 | |
uint16 | |
uint32 | |
uint64 |
| Minimum Value | Maximum Value | Resulting Type |
|---|---|---|
| -2147483648 | 2147483647 | int32 |
| 2147483648 | 4294967295 | uint32 |
| 4294967296 | 9223372036854775807 | int64 |
| -9223372036854775808 | -2147483649 | |
| 9223372036854775808 | 18446744073709551615 | uint64 |
6.8.3.2.Floating Point Constants🔗
Syntax
Constraints
- An unsuffixed floating point constant must contain a valid floating point value that can be represented as a 64-bit floating point value.
- A suffixed floating point constant must contain a valid floating value that can be represented by the type specified by the suffix. (See Table 11 - Floating point suffix type table.)
- Hexadecimal/Octal/Binary floating point numbers, are explicitly not supported in Tetral grammar.
Semantics
- The result of an unsuffixed floating point constant is a 64-bit floating point number.
- The result of a suffixed floating point constant is a floating point value of the type matching the suffix. (See Table 11 - Floating point suffix type table.)
| Suffixes | Type |
|---|---|
float32 | |
float64 |
6.8.3.3.Character Constants🔗
Syntax
Constraints
- A character constant's contained codepoint cannot be larger than one which can be encoded in a single UTF-8 byte.
Semantics
- Character constants result in a single unsigned 8-bit integer, equal to the character constant's UTF-8 encoded byte value.
- An escaped sequence leads to the escaped character's value being used. See the table below for escaped character values.
| Escape Sequence | Resulting codepoint name. | Unicode codepoint. |
|---|---|---|
| Single quote character | U+0027 | |
| Double quote character | U+0022 | |
| Slash character | U+002F | |
| Tabulation character | U+0009 | |
| Linefeed character | U+000A | |
| Carriage Return character | U+000D | |
| Dollar Sign character | U+0024 |
6.8.3.4.Predefined Constants🔗
Syntax
Semantics
- The result of the
true andfalse keywords are a boolean value of the corresponding value. - The result of the
null value is a pointer with a null value. A null value can be implicitly cast to any nullable reference type or any nullable value-type-reference type. Thenull value can also be assigned to any type marked as nullable with the? keyword. - The result of the
nan keyword is a 32-bit floating point value with the special Not-A-Number value. - The result of the
inf ,-inf and+inf keywords are a 32-bit floating point value with an infinite value of the corresponding signedness. If no sign is specified, then the result the same as with the+ sign.
6.8.4.Operators🔗
Syntax
6.8.5.Semicolons🔗
Syntax
Lexical Analysis Rules
- The semicolon as a marker for the end of an expression or for a statement is completely optional and need not be used unless there may be ambiguity in the syntax.
- Certain syntactic elements (For example, the
for-loop-statement ) explicitly use semicolons to delimit parts of the grammar. In such cases the semicolon is required.
6.8.6.Comments🔗
Except within a string or character literal, the characters
Except within a string or character literal, the characters
Except within a string or character literal, the characters
In almost all cases, comments are simply skipped over and ignored, except for
documentation comments (Starting with
6.9.Expressions🔗
6.9.1.Primary Expressions🔗
Syntax
6.9.1.1.New Expression🔗
Syntax
Constraints
- New expressions can only be used to instantiate structs, classes and arrays.
Semantics
- Instantiates a new instance of a specific type.
- If the type name was not specified, the type must be inferred from the surrounding context.
Object Instantiation
Constraints
- The specified arguments must match a constructor that exists on the type.
Semantics
- Instantiates an object by calling the appropriate constructor.
Array Instantiation
Constraints
- The specified expression must result in an unsigned 32-bit integer, this expression's result will be referred to as the instantiated array's length.
- The array type must be default initialize-able.
Semantics
- Instantiates an array of the defined length and initializes each value in the created array to the array type's default initialized value.
- If the specified length of the array is 0, then the array is not allocated and is instead assigned to a null pointer, aka, an empty array.
6.9.1.2.'Super' Expression🔗
Syntax
Constraints
- May only be used inside a class or interface method declaration or a class constructor declaration.
Semantics
- The result of the expression is a pointer to the
this object's parent object in the current context.
6.9.1.3.'This' Expression🔗
Syntax
Constraints
- May only be used inside a class, struct or interface method declaration or a class constructor declaration.
Semantics
- The result of the expression is a reference to the
this object in the semantic context in which it used.
6.9.1.4.Cast Expression🔗
Syntax
Constraints
- Cast expressions can only cast between types which have a defined implicit or explicit conversion. (See 4.14. Type Conversions.)
- Bit-casting can only be performed on arithmetic types.
- Bit-casting can only cast from one type to another type of the same size.
Semantics
- Casts a value returned by an expression to a different type.
- Regular casting operations are defined in 4.14. Type Conversions.
- Bit-casing only treats the expression's value as a different type, no conversion operations are performed.
6.9.1.5.Object Literals🔗
Syntax
Constraints
- If an object literal's struct type cannot be inferred from the surrounding context, then the expression is invalid.
- An object literal may only initialize properties that exist on the struct type.
- An object literal may not declare any duplicate properties.
Semantics
- The type of struct initialized by an object literal must be inferred from the surrounding context.
- Semantically, an object literal is identical to constructing the struct type and setting each property separately.
6.9.1.6.Array Literals🔗
Syntax
6.9.1.7.String Literals🔗
Syntax
Constraints
- All expressions and identifiers declared inside a string must return a non-void value.
Semantics
- A string literal will always result in a value of a
stringtype. - It can be assumed that each line in a multiline string will have some level of indentation. After parsing, this indentation must be stripped from the start of each line. Since indentation can vary from line to line, the amount indentation removed depends on the line with the least indentation.
- The template expressions
$ identifier and${ expression } with an identifier as the expression are semantically identical. - For templated strings, each templated expression is evaluated in declaration order and appended to the resulting string. (See 6.6. String Conversion Semantics.)
- Semantically, a templated string literal is identical to declaring each part of the string as its own separate expression and concatenating it together.
6.9.2.Trail Expressions🔗
Syntax
6.9.3.Index Access Expression🔗
Syntax
Constraints
- The type returned by the target
trail-expression must be an indexable type. (See Table 13 - Type Indexing) - The indexing expression must be a type used to index into the target expression's type. (See Table 13 - Type Indexing)
- If the type being accessed is a
string, then the index may not be written to.
Semantics
- Accesses an indexable element in an indexable value.
- A valid index access expression is a valid lvalue.
- Accessing a non-readonly array is a valid mutable lvalue.
Runtime Errors
- It is a runtime error if the index being accessed is beyond the length of the indexed operand.
| Type | Indexing Type | Value type returned by index access |
|---|---|---|
| String Type | uint32 | Single UTF-8 byte, (uint8) |
| Any Array Type | uint32 | The array's component type |
6.9.4.Property Access Expression🔗
Syntax
Constraints
- The property being accessed must exist on the object operand.
- A property access must not violate class/interface access semantics.
- The
. operator cannot be used on the object operand. Either the!. or?. operator must be used by the user to acknowledge the nullability value.
Semantics
- The result of the
. and!. operator is a copy of the value of the property or virtual property being accessed. - The result of the
?. operator is a nullable copy of the value of the property or virtual property being accessed. If the object operand is an empty nullable, then the resulting value is an empty nullable. - Accessing a property of a nullable type with the
?. operator causes the nullable to be flattened. - A valid property access expression is a valid lvalue.
Runtime Errors
- It is a runtime error if the object operand is an empty nullable value.
6.9.5.Function Call Expression🔗
Syntax
Constraints
- The value returned by the target trail-expression must be a function.
- All arguments passed to the function must be assignable to that function's declared arguments.
Semantics
- Invokes a function with the specified arguments.
- Calling arguments are evaluated from left to right.
- The value returned by the expression is dependent on the return type of the function being invoked.
- When the
trail-expression matches multiple overloading functions, the one that would require the least number of type conversions is selected.
6.9.6.Coalescence Expression🔗
Syntax
Constraints
- The left operand must be a nullable value.
- The right operand must be assignable to the left operand, or be a non-nullable value of an assignable type to the left operand.
Semantics
- The result of the
?? operator is the left operand if it is present. Otherwise, the result of the operator is the right operand. - If the left operand is returned, the right operand must not be evaluated.
6.9.7.Unary Expressions🔗
Syntax
6.9.7.1.Unary Increment Expressions🔗
Syntax
Constraints
- The operand must be a valid lvalue that can be modified.
- The operand must be of an integral or floating point type.
Semantics
- If the increment operator is placed before the operand, the result is the value of the operand incremented by 1. As a side effect, the resulting value is written to the operand object.
- If the increment operator is placed after the operand, the result is the value of the operand. As a side effect, the operand's value is incremented by 1 and written to the operand object.
6.9.7.2.Unary Decrement Expressions🔗
Syntax
Constraints
- The operand must be a valid lvalue that can be modified.
- The operand must be of an integral or floating point type.
Semantics
- If the decrement operator is placed before the operand, the result is the value of the operand decremented by 1. As a side effect, the resulting value is written to the operand object.
- If the decrement operator is placed after the operand, the result is the value of the operand. As a side effect, the operand's value is decremented by 1 and written to the operand object.
6.9.7.3.Unary Arithmetic Expressions🔗
Syntax
Constraints
- For the
- operator, the operand must be of an arithmetic type. - For the
+ operator, the operand must be of an arithmetic type or a boolean. (Not boolean-like, an explicit boolean.)
Semantics
- The result of the - operator is the negative of the operand's value, or positive, if the operand's value is already negative. If the operand's type an unsigned type, the resulting type is a signed type of the same size.
- The result of the + operator, for arithmetic values, is the value of the operand.
- The result of the + operator, for boolean values, is the value 0, if the boolean value is false, otherwise the result is 1.
6.9.7.4.Unary Bitwise Negation Expressions🔗
Syntax
Constraints
- The operand must be a valid lvalue.
- The operand must be of an integral type.
Semantics
- The result of the expression is the value of the operand with all the bits flipped (All zeroes set to ones and vice versa)
6.9.7.5.Unary Logical Negation Expressions🔗
Syntax
Constraints
- The operand's value must be boolean compatible. (Defined in 4.14.3. Boolean Conversion.)
Semantics
- The result of the
! operator is the value of the operand, converted to a boolean, and then the value inverted.
6.9.8.Power (POW) Expression🔗
Syntax
Constraints
- Both operands must be of a compatible arithmetic type.
Semantics
- The usual arithmetic conversions are applied.
- The result of the ** operator is the left operand raised to the power of the right operand.
6.9.9.Multiplicative Expressions🔗
Syntax
Constraints
- Both operands must be of a compatible arithmetic type.
- In the case of the
* operator, if the left operator is a string, the right operand must be an integral.
Semantics
- The usual arithmetic conversions are applied.
- The result of the
* operator, when applied to two arithmetic types, is the left operand multiplied by the right operand. - The result of the
* operator, when applies to a string operand and an integral operand, is the value of the string operand with its content repeated an amount of times specified by the integral operand. - The result of the
/ operator is the left operand divided by the right operand. - The result of the
% operator is the remainder of dividing the left operand by the right operand.
6.9.10.Additive Expressions🔗
Syntax
Constraints
- For the
- operator, both operands must be a compatible arithmetic type. - For the
+ operator, either both operands must be of a compatible arithmetic type, or the left operand must be a string, and the other operand must be a non-void value.
Semantics
- The usual arithmetic conversions are applied.
- The result of the
+ operator, when both operands are arithmetic, is the left operand plus the right operand. - The result of the
+ operator, when the left operand is a string, is the string operand's string value concatenated with the string representation of the other operand. (See 6.6. String Conversion Semantics.) - The result of the
- operator is the left operand minus the right operand.
6.9.11.Shift Expressions🔗
Syntax
Constraints
- Both operands must be of an integral type.
Semantics
- The result of the << operator is the value of the left operand with its bits shifted to the left by the value of the right operand.
- The result of the >> operator is the value of the left operand with its bits shifted to the right by the value of the right operand.
- In the case the right operand is a negative value, the operation acts as the inverse operation: left shift becomes right shift and vice versa.
6.9.12.Relational Expressions🔗
Syntax
Constraints
- Both operands must be of a compatible comparable type (See 6.2. Comparison Semantics.)
Semantics
- If both operands are compatible arithmetic types, then the usual arithmetic conversions are applied.
- The result of the < operator is a boolean true value if the left operand's value is less than the right operand's value.
- The result of the <= operator is analogous to the result of the < operator, except that it also returns a boolean true value when the left operand's value is equal to the right operand's value.
- The result of the > operator is a boolean true value if the left operand's value is greater than the right operand's value.
- The result of the >= operator is analogous to the result of the > operator, except that it also returns a boolean true value when the left operand's value is equal to the right operand's value.
6.9.13.Equality Expressions🔗
Syntax
Constraints
- Both operands must be of a compatible type.
Semantics
- If both operands are compatible arithmetic types, then the usual arithmetic conversions are applied.
- The result of the == operator is a boolean true value if the left operand's value is equal to the right operand's value.
- The result of the != operator is the inverse of the == operator's value.
6.9.14.XOR Expressions🔗
Syntax
Constraints
- Both operands must be of a valid integral type
Semantics
- The usual arithmetic conversions are applied.
- The result of the
^operator is the value of the left operand with every bit set to true if and only if the corresponding bit is set in only one of the operands.
6.9.15.Bitwise AND Expressions🔗
Syntax
Constraints
- Both operands must be of a valid integral type
Semantics
- The usual arithmetic conversions are applied.
- The result of the
&operator is the value of the left operand with every bit set to true if the corresponding bit in the right operand is also set to true.
6.9.16.Bitwise OR Expressions🔗
Syntax
Constraints
- Both operands must be of a valid integral type
Semantics
- The usual arithmetic conversions are applied.
- The result of the
|operator is the value of the left operand with every bit set to true if the bit is set in either the left operand or the right operand.
6.9.17.Logical AND Expressions🔗
Syntax
Constraints
- Both operands must be of a valid integral or boolean-like type.
Semantics
- The usual arithmetic conversions are applied if both operands are arithmetic types.
- The result of the
&&operator is a boolean true value if both operands evaluate to a true value. - If the first operand evaluates to a false value, the second operator must not be evaluated.
6.9.18.Logical OR Expressions🔗
Syntax
Constraints
- Both operands must be of a valid integral or boolean-like type.
Semantics
- The usual arithmetic conversions are applied if both operands are arithmetic types.
- The result of the
||operator is a boolean true value if at least one of the operands are a boolean true value. - If the left operand evaluates to a true value, the right operand must not be evaluated.
6.9.19.Conditional Expression🔗
Syntax
Constraints
- The condition expression must result in a boolean-like value.
- Both the left and the right operands must result in a compatible type.
Semantics
- The condition operand is evaluated first, if it results in a true value, the left operand is the result of the expression, otherwise the right operand is the result.
- If the left operand is the result, the right operand must not be evaluated, and vice versa.
6.9.20.Assignment Expression🔗
Syntax
6.9.21.Top Level Expression🔗
Syntax
6.9.22.Loop Condition Expression🔗
Syntax
Constraints
- Loop condition expressions must always result in a value that can be assigned in some way to a true/false value.
Description
Loop condition expressions declare an expression that may or may not be surrounded by parentheses. Loop condition expressions declared with parentheses cannot have any part of the expression outside the parentheses
6.10.Type Expressions🔗
Syntax
6.10.1.Const Type Expressions🔗
Syntax
6.10.2.Type Name Expressions🔗
Syntax
6.10.3.Parameterized Type Expressions🔗
Syntax
6.10.4.Nullable Type Expressions🔗
Syntax
Semantics
- Declares a type to be nullable.
- Variables, properties, fields or function arguments marked as nullable may
be set as
null .
6.10.5.Array Type Expressions🔗
Syntax
6.11.Declarations🔗
Syntax
6.11.1.Declaration Modifiers🔗
Syntax
Constraints
- No declaration modifier may be set more than once.
6.11.2.Type Parameter Lists🔗
Syntax
6.11.3.Module Declarations🔗
Syntax
Concepts
- Module name
- The name of the module itself declared after the
modulekeyword. - Native Library Name
- The string literal declared after the
fromkeyword in a native module declaration. It is a file name for the dynamically-loadable native C/C++ library that any native functions declared in the file will be linked to. The string must not include the file extension. (eg:.dll)
Constraints
- The file name/path referenced by the native library path, if present, must be a valid path to a dynamically linkable library. The path must be accessible to the compiler and known at compile time.
- It is not required for a module declaration to declare a name unique to the file it's declared in. Multiple files may share the same namespace.
Semantics
- A module declaration is used by import statements to know from which file to import certain symbols.
- Adding a module declaration is fully optional.
- Omitting the module declaration means the file will be able to export any symbols.
6.11.4.Import Declarations🔗
Syntax
Constraints
- The import statement must result in at least one symbol being included in the source file. If no symbols are found with the namespace specified in the import, then the import is considered a failure.
Semantics
- Import all export symbols with the specified module name namespace into the source file.
- All IR, or source files, the compiler is able to find that share the specified namespace are imported into the current source file.
- When an IR file, or valid source file is found that is declared as native, the native module need not be loaded.
- During importing, any currently existing native bindings should also be checked for import. This is because native bindings can be registered without a pre-existing Tetral IR or source file to represent the declarations.
6.11.5.Function Declarations🔗
Syntax
Constraints
- The name and parameter types must not match a function that already exists in the current scope.
- Function argument lists can only ever have one variadic argument: the last one. Multiple variadic arguments or a variadic argument not being the last one are invalid.
- Functions declared with the
exportkeyword can only be declared in the global scope. - Functions declared with the
nativekeyword must not have a function body.
Description
Declares a function that will be accessible in the current scope.
6.11.6.Extension Function Declarations🔗
Syntax
Constraints
- The type referenced in the
extension-function-name must be a valid type. - The function may not access any private variables declared in the extended type if it is a class type.
Semantics
- Declares a function which can be called and accessed as though it were a
method belonging to any arbitrary type.
6.11.7.Variable Declarations🔗
Syntax
Constraints
- Must not have the name of an already accessible variable in the current scope.
- If a value is specified for the declaration, the variable's resulting value must match the declared type, or be easily convertible to the variable's type.
- If the variable is declared with
const , then a value must be specified. - Variable declared with the
const modifier may not be modified at a later time. - Variables declared with the
export keyword must only be declared in the global scope. - Variables cannot be declared with the following keywords:
native
- If an initializer value is present, the initializer cannot reference the declared variable.
- If the variable's type cannot be default-initialized, then a value must be explicitly specified. (See 6.5. Default Initialization)
Semantics
- Declares a variable or a constant that will be accessible in the current scope.
- If an initializer value is present, it becomes the value of the variable.
- Variables declared with the
export keyword are made available for access outside the declaring file.
If no initializer value is set, then the value must be default-initialized. (See 6.5. Default Initialization)
6.11.8.Struct Declarations🔗
Syntax
Constraints
- A struct must not be declared with the
native modifier. - Struct's must only be declared in the main scope of a source file.
- Each struct property must have a unique name.
- Each struct property, if a default value is set, must have an expression whose type is assignable to the property's type.
- Struct properties may not directly or indirectly reference the declared struct type unless through a nullable value-type-reference.
Semantics
- Declares a struct data type and its properties.
6.11.9.Enum Declarations🔗
Syntax
Constraints
- An enum type may not be declared with the same name as an existing interface, enum, struct, class or function in the same namespace.
- Each enum constant declaration must have a unique name.
Semantics
6.11.10.Variant Declarations🔗
Syntax
Constraints
- A Variant may not be declared with the same name as an existing declaration in the same namespace.
- Each declared variant of a variant type declaration must have a unique name.
Semantics
- Declares a variant type with the specified name and variants.
- Declares each variant's properties.
6.11.11.Interface Declarations🔗
Syntax
Constraints
- Interfaces can only be declared in a source file's global scope.
Semantics
6.11.12.Class Declarations🔗
Syntax
Constraints
- Classes may only be declared in a file's global scope.
- Classes must not be declared with the same name as an already existing function, struct, enum, interface or global variable.
- Classes implementing interfaces must declare a valid method implementation for every function outlined by its interfaces.
- Classes may not implement 2 or more interfaces with methods whose implementation would cause other constraints to be violated.
- Classes may not extend themselves, directly or indirectly.
- Classes may not extend a class that has not been marked with the
extendable keyword. - Classes may only extend other valid classes.
- Classes may only implement valid interfaces.
- Classes may not extend classes if the extension would cause an inheritence cycle.
Semantics
- The same padding and alignment semantics apply to classes as do to structs.
6.11.12.1.Class Fields🔗
Constraints
- Classes may not have 2 or more fields with the same name.
- Fields declared with the
constmodifier must either have a value specified by default or be initialized in the constructor. - If a field has a default value, its value must be a valid expression that results in a type compatible with the field's type.
Semantics
- Declares an accessible property.
6.11.12.2.Class Constructors🔗
Constraints
- Classes may not have 2 or more constructors with the same function signature.
- Constructors may not be declared with the
overrideorstatickeyword. - Constructors may not return a value.
- All constraints that apply to regular function argument lists also apply to constructor arguments.
- If a class is declared with a base type that requires explicit
initialization, then each constructor declared in a class must
initialize the base type with a
super-constructor-expression .
Semantics
6.11.12.3.Class Methods🔗
Constraints
- The declared return type must be a valid type.
- The same constraints that apply to regular function declarations also apply to method declarations, except for those which cannot be applied to method declarations.
- Classes may only declare functions with the same name if the argument signatures differs.
- Classes may not override any methods declared in an inherited type that have
the
const keyword. - Classes may not override any methods declared in an inherited type that have
the
constkeyword. - Classes overriding methods from an inherited class or interface must use the
overridemodifier. - A method declared with the
static modifier cannot also use the modifieroverride .
Semantics
- Declares a member function that belongs to the class the method was declared in.
6.11.12.4.Class Access Modifiers🔗
Constraints
- Each modifier is mutually exclusive to another access modifier.
Semantics
- Access modifiers modify how fields and methods may be accessed from outside the class the symbol was declared in.
- The
private keyword ensures a class symbol is only accessible in the class it was declared in. - The
protected keyword ensures a class symbol is only accessible in the class it was declared in, or in a class extending the class. - The
public modifier allows a class symbol to be accessed from anywhere. - Omitting the access modifier is equivalent to using the
public access modifier.
6.12.Statements🔗
Syntax
6.12.1.Block Statements🔗
Syntax
Description
Declares a list of statements delimited by { and
} characters.
Variables declared inside of block statements cannot be referenced from outside of that block.
6.12.2.Loop Flow Control Statements🔗
Syntax
Description
Influences the way a loop executes. Continue statements instruct the loop to move on to the next iteration of execution, while break statements cause the loop to be exited early before its condition can evaluate to a false value.
Control Flow statements can be declared with a label. This label is optional, if no label is specified, the statement changes how the immediate loop behaves. If a label is set, then it controls how the referenced loop behaves.
Constraints
- Can only be used inside a loop
- If a label is specified, the label can only reference a label of a loop the statement is inside.
6.12.3.If Statements🔗
Syntax
Description
Evaluates a condition and executes either a body, an else statement or moves on without executing.
Constraints
- The if statement's condition must result in a value that can be assigned to a true/false value.
6.12.4.Labeled Statements🔗
Syntax
6.12.5.For Loop Statements🔗
Syntax
Terminology
A for loop has three operands, the first, second and third operand. The first, if present, declares a variable, the second evaluates a condition, and a third expression performs some operation to move the for-loop forward.
Constraints
- The second operand, if present, must result in a boolean-like value.
Semantics
6.12.6.For-Each Loop Statement🔗
Syntax
Constraints
- The right operand, after the
identifier , must result in an array or a string. - The variable declaration must be of the type that the right operand contains.
Semantics
- Iterates over each element of the right operand and performs some operation.
- The element being iterated over is the one declared in the
identifier and must be available as a variable.
6.12.7.Do While Loop Statements🔗
Syntax
Constraints
- The loop's condition must result in a value that can be evaluated as a boolean.
Description
Evaluates a block of code until the defined loop condition results in a false value. The loop's block will be evaluated at least once until the condition is reached.
6.12.8.Regular While Loop Statements🔗
Syntax
Constraints
- The loop's condition must result in a value that can be evaluated as a boolean.
Description
Evaluates a block of code until the defined loop condition results in a false value. If the condition results in a false value when the loop is first reached, the loop's block is not evaluated at all.
6.12.9.Return Statements🔗
Syntax
Description
Causes the current function to halt execution and return to the caller.
Constraints
- Can only be used inside a function body.
- If the return statement is used in a non-void function, it must always return a value.
- If the return statement is used in a void function, it must never return a value.
- The value returned by a return statement, must be assignable to the function's return type.
6.12.10.Assert Statements🔗
Syntax
Semantics
- Evaluates the condition operand. If the condition results in a true boolean-like value, nothing happens. Otherwise, an error is thrown.
- If the message operand is present, it is evaluated and used as the error message for the thrown error.
- If no message operand is present, it must be created from the values of the variables present in the condition expression. (See Message Generation section below).
- The inclusion of the assert statement in the compiled output is optional, as is the execution of the assert statement
Constraints
- The expression the assert statement evaluates must always result in a Boolean-like value.
- The message expression, if present, must always result in a String value.
Message Generation
The messages used by assert statements, if one is not explicitly provided, should be composed of the values of each variable used in the expression, for example, when asserting that a variable named "foo" is equal to 10, the automatic error message should be composed of the expression string and followed by the variable's value. In effect, turning the statement into this:
| |
Examples
| |
6.13.Translation Unit🔗
7.Intermediate Representation (Bytecode Format)🔗
The result of the compilation of a Tetral source file should be a Tetral Bytecode file. This section outlines the bytecode format.
7.1.File Header🔗
Each Tetral IR file is prefixed with a file header which provides information about the file's contents. The layout of the header is as follows.
The header begins with an ASCII string that must always have the value
TetralLang. This is immediately followed by a 2 byte unsigned
integer denoting the file's version.
The file version integer is an incrementing integer which specifies what version of the Tetral bytecode format the file complies with.
| Version number | Description |
|---|---|
| 1 | The first version of the Tetral Language |
The vile version integer is then followed by a repeating pattern. Each instance of this pattern begins with a single unsigned byte, the entry type. No entry may be defined twice, but not all entries are required to appear in a file.
The following types are valid.
| Name | Hexadecimal Value | Description |
|---|---|---|
END | 0x00 | Special entry type used to mark the end of the header. |
STRTABLE | 0x01 | String table information |
STRDATA | 0x02 | String pool information |
TYPETABLE | 0x03 | Type table information |
FUNCTABLE | 0x04 | Function table information |
INSTRBLOCK | 0x05 | Instruction block information |
RDATA | 0x06 | Read only data block information |
MDATA | 0x06 | Pre-defined metadata properties |
ADATA | 0x06 | List of arbitrary key,value string pairs. |
MODINFO | 0x07 | Information about the import and export list. |
7.1.1.The STRTABLE Entry🔗
Comprised of a 64-bit unsigned integer: an offset of the string table's data from the start of the file, and a 32-bit unsigned integer: the number of entries in the string table.
(See 7.2. String Pool Table)
7.1.2.The STRDATA Entry🔗
Comprised of two 64-bit unsigned integers, the first is an offset and the second is a size. The offset specifies where, from the start of the file, the file's string pool data is located and the size specifies the size (in bytes) of the file's string pool data.
(See 7.3. String Pool Data)
7.1.3.The TYPETABLE Entry🔗
Comprised of a 64-bit unsigned integer: an offset of the type table's data from the start of the file, and a 32-bit unsigned integer: the number of entries in the type table.
(See 7.4. Type Table)
7.1.4.The FUNCTABLE Entry🔗
Comprised of a 64-bit unsigned integer: an offset of the function table's data from the start of the file, and a 32-bit unsigned integer: the number of entries in the function table.
(See 7.5. Function Table)
7.1.5.The INSTRBLOCK Entry🔗
Comprised of the following orders. The values appear in the binary in the order they are defined here.
- Instruction block offset
- A 64-bit unsigned integer that specifies the file's instruction block offset from the start of the file.
- Instruction block size
- A 32-bit unsigned integer that specifies the file's instruction block size in bytes.
- Instruction count
- A 32-bit unsigned integer that specifies how many instructions are in the instruction block.
(See 7.6. Instructions Block)
7.1.6.The RDATA Entry🔗
Comprised of two 64-bit unsigned integers, they are as follows and appear in the order they are listed.
- The offset of the data block from the start of the file.
- The size of the data block, in bytes.
(See 7.7. Data Block)
7.1.7.The MDATA Entry🔗
Comprised of the following values which and appear in the order they are listed.
- The file's global scope size
- A 64-bit unsigned integer that determines how many bytes are required to represent the variables stored in the file's global scope.
- Entrypoint function index
An 32-bit unsigned integer index into the function table from where the execution of the instructions in the file should begin.
A value (in hexadecimal) of
0xFFFFFFFFmeans that the file has no entrypoint and cannot be executed by itself.
7.1.8.The ADATA Entry🔗
This entry begins with a 32-bit unsigned integer that denotes the number of key value pairs stored in the section.
It is then followed by that many number of key-value pairs of strings. Each string is prefixed with a 16-bit unsigned integer denoting the length of the string and is then followed by that many bytes of UTF-8 bytes.
7.2.String Pool Table🔗
The String Pool Table, or just the String Table, is a list of 32-bit unsigned
integer pairs. The number of entries in the table is determined in the
STRTABLE entry in the header.
Each 32-big unsigned integer is defined as follows:
- String Pool Offset
- The first 32-bit unsigned integer is an address of a string's content in the String Pool Data.
- String Length
- The second integer is the number of UTF-8 code units in the string.
The first entry in the list may be a "null" string entry, meaning an entry with offset 0 and length 0, used as a placeholder for an empty, or null, string.
If this section is present in a bytecode file, then the String Pool Data section must be present too, as well as the corresponding entries in the header.
7.3.String Pool Data🔗
A block of UTF-8 encoded codepoints. The size and location of this section is
determined by the STRPOOL entry in the header.
If this section is present in a bytecode file, then the String Pool Table section must be present too, as well as the corresponding entries in the header.
7.4.Type Table🔗
This section is a list of type definitions. The number of definitions in this
section and the section's location in the file is determined by the
TYPETABLE entry in the header.
Each definition is prefixed with 2 values, an unsigned 32-bit integer and an 8-bit unsigned integer.
The first integer is the "type index" of the type definition. This index is used in function table entries, other type table entries, and in some IR instructions to reference the type definition.
The second 8-bit unsigned integer marks the type of the entry. The following values are considered valid.
| Name | Hexadecimal Value |
|---|---|
ARRAY | 0x00 |
FUNCSIGN | 0x01 |
STRUCT | 0x02 |
ENUM | 0x03 |
INTERFACE | 0x04 |
CLASS | 0x05 |
NULLABLE | 0x06 |
PARAMETERIZED | 0x07 |
TYPEPARAMETER | 0x08 |
PARAMLIST | 0x09 |
IMPORTED | 0x0A |
Depending on the type, one of the schemas defined in the following sections is used for the type.
Some entries may reference a "null type index," this is a type index
constant with the hexadecimal value of 0xFFFFFFFF and means that
the type index references no type
7.4.1.Arrays, Nullable Definitions and Reference Definitions🔗
These are defined in one section since their schema is identical, but each type results in a different type construct.
These entries define only a single 32-bit unsigned integer value. It is the type index of the type they reference.
The referenced type index may not be a null type index.
7.4.2.Function Signatures🔗
Function signatures are defined with the following values. The values appear in the binary in the order they are defined here.
- Return type index
- 32-bit unsigned integer index into the type table, referencing the type the function signature uses as a return type. This may not be a null type index.
- Variadic state
- 8-bit unsigned integer with a value of either 0 or 1 used to determine if the function takes in a variable number of arguments.
- Parameter list index
- A 32-bit unsigned integer referencing the
PARAMLISTentry that outlines the type arguments for this type. If the type has no type arguments, this will be a null type index. - Arguments count
- A 32-bit unsigned integer stating how many entries will be in the next argument list.
- Argument type indexes list
An array of 32-bit unsigned type indexes referencing the type of the corresponding function argument's type.
None of these may be null type indices. If the variadic state of the function is set to 1, then the last parameter to appear in this list must reference an entry with the type
ARRAY.
7.4.3.Structs🔗
Structs are defined with the following values. The values appear in the binary in the order they are defined here.
- Name index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's name.
- Namespace index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's namespace. If the type has no namespace, the referenced entry will be the table's null entry.
- Type parameter list index
- A 32-bit unsigned integer referencing the
PARAMLISTentry that outlines the type arguments for this type. If the type has no type arguments, this will be a null type index. - Flags
- Property count
- Property list
7.4.4.Enums🔗
Enumerations are defined with the following values. The values appear in the binary in the order they are defined here.
- Name index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's name.
- Namespace index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's namespace. If the type has no namespace, the referenced entry will be the table's null entry.
- Type parameter list index
- A 32-bit unsigned integer referencing the
PARAMLISTentry that outlines the type arguments for this type. If the type has no type arguments, this will be a null type index. - Constant count
- Constant list
7.4.5.Interfaces🔗
Interfaces are defined with the following values. The values appear in the binary in the order they are defined here.
- Name index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's name.
- Namespace index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's namespace. If the type has no namespace, the referenced entry will be the table's null entry.
- Type parameter list index
- A 32-bit unsigned integer referencing the
PARAMLISTentry that outlines the type arguments for this type. If the type has no type arguments, this will be a null type index. - Extends count
- Extends list
- Member count
- Method list
7.4.6.Classes🔗
Classes are defined with the following values. The values appear in the binary in the order they are defined here.
- Name index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's name.
- Namespace index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's namespace. If the type has no namespace, the referenced entry will be the table's null entry.
- Type parameter list index
- A 32-bit unsigned integer referencing the
PARAMLISTentry that outlines the type arguments for this type. If the type has no type arguments, this will be a null type index. - Parent type index
- Implements count
- Implements list
- Member count
- Method list
7.4.7.Parameter List Instantiation🔗
Parameter lists define the types another type requires to be complete.
Parameter lists are defined with the following values. The values appear in the binary in the order they are defined here.
- Referenced type index
- Type count
- Type list
7.4.8.Imported Type🔗
Imported types are types which are not defined in the current file, but are instead defined in another file and must be imported. They are defined with the following values. The values appear in the binary in the order they are defined here.
- Name index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's name.
- Namespace index
- 32-bit unsigned integer index referring to the entry in the file's string table that is used as the type's namespace. If the type has no namespace, the referenced entry will be the table's null entry.
7.5.Function Table🔗
7.6.Instructions Block🔗
7.7.Data Block🔗
7.8.Import List🔗
7.9.Export Table🔗
8.Bytecode Instructions🔗
8.1.General Purpose Instructions🔗
8.2.Stack Memory Instructions🔗
8.3.Global Memory Instructions🔗
8.4.Closure Instructions🔗
8.5.Heap Memory Instructions🔗
8.6.Function Call Instructions🔗
8.7.Foreign Symbol Instructions🔗
8.8.Conversion Instructions🔗
8.9.Unary Instructions🔗
8.10.Binary Instructions🔗
8.10.1.Integer-only Binary Instructions🔗
8.10.2.Boolean-only Binary Instructions🔗
8.10.3.Arithmetic Binary Instructions🔗
8.10.4.Comparison Instructions🔗
8.10.5.String/Array Instructions🔗
9.Parsing and Compilation🔗
10.Interpreter🔗
11.The Standard Library🔗
This section defines the modules that are part of the Tetral Standard Library. Whether any function defined in the following sections is implemented natively or within Tetral itself is explicitly left up to implementations to decide on.
11.1.The std Module🔗
The most basic standard library functions that would be required by a program.
11.1.1.The Optional Variant🔗
variant Optional <T > {
EMPTY,
PRESENT(T value)
}The optional type is monadic type to represent a value which either is present or empty.
11.1.2.The Result Variant🔗
variant Result <V , E > {
ERROR(E err),
SUCCESS(V value)
}11.1.3.The readln Function🔗
string readln ()Asks for input from the standard input and blocks execution until input has been given.
11.1.4.The println Function🔗
void println (string message)Prints a string to the standard output and appends a newline (U+000A) character.
11.1.5.The getenv Function🔗
string getenv (string name)The getenv function takes as input a single string: the
variable name and returns a value associated with that environment variable.
If no value is associated with that value, a null/empty string is returned instead.
11.1.6.The setenv Function🔗
void setenv (string name, string value)
void setenv (string name, string value, bool overwrite)This function is defined with two overloads. Both definitions take in a variable name and value. One of the overloads takes in a boolean value. If this value is false, then the envirnoment variable will not be overwritten if it is already set.
The function changes a value in or adds to the list of environment variables.
11.1.7.Numeric Parsing Functions🔗
int32 ? parseInt (string s)
uint32 ? parseUint (string s)
int64 ? parseInt64 (string s)
uint64 ? parseUint64 (string s)
float32 ? parseFloat (string s)
float64 ? parseFloat64 (string s)Every one of these functions takes a string as an input and parses the beginning of that string into a number of the equivalent value of the respective type.
11.2.The std::strings Module🔗
string stringFromUtf8 (uint8 [] utf8bytes)
string stringFromUtf16 (uint16 [] utf16bytes)
string stringFromUtf32 (uint32 [] codepoints)
string string ::trim ()
string string ::trimLeft ()
string string ::trimRight ()
string string ::substring (uint32 start)
string string ::substring (uint32 start, uint32 end)
uint32 ? string ::indexOf (uint8 utf8char)
uint32 ? string ::indexOf (uint8 utf8char, uint32 start)
uint32 ? string ::indexOf (uint32 codepoint)
uint32 ? string ::indexOf (uint32 codepoint, uint32 start)
uint32 ? string ::indexOf (string other)
uint32 ? string ::indexOf (string other, uint32 start)
uint32 ? string ::lastIndexOf (uint8 utf8char)
uint32 ? string ::lastIndexOf (uint8 utf8char, uint32 end)
uint32 ? string ::lastIndexOf (uint32 codepoint)
uint32 ? string ::lastIndexOf (uint32 codepoint, uint32 end)
uint32 ? string ::lastIndexOf (string other)
uint32 ? string ::lastIndexOf (string other, uint32 end)
bool string ::contains (uint8 utf8char)
bool string ::contains (uint8 utf8char, uint32 end)
bool string ::contains (uint32 codepoint)
bool string ::contains (uint32 codepoint, uint32 end)
bool string ::contains (string other)
bool string ::contains (string other, uint32 end)
bool string ::replace (string sequence, string with)
bool string ::startsWith (string other)
bool string ::endsWith (string other)
uint32 ? string ::codepointAt (uint32 index)
uint32 string ::codepoints ()
uint32 [] string ::toUtf32 ()
uint16 [] string ::toUtf16 ()
uint8 [] string ::toUtf8 ()
bool string ::isEmpty ()
bool string ::isBlank ()
string string ::toUpperCase ()
string string ::toLowerCase ()11.3.The std::math Module🔗
11.4.The std::io Module🔗
11.5.The std::reflection Module🔗
11.6.The std::json Module🔗
12.Compiler Command Line Interface🔗
This section details the CLI (Command Line Interface) of the default Tetral implementation.
The usage of the tetral command looks like so:
tetral [OPTIONS] [COMMAND] [COMMAND ARGUMENTS]
12.1.Commands🔗
12.1.1.The compile Command🔗
Command Usage
tetral [OPTIONS] compile [IN FILES]
Description
The compile command takes in a list of file paths that are parsed and then compiled into one single bytecode format. The files in the list can be source files or bytecode files.
12.1.2.The run Command🔗
Command Usage
tetral [OPTIONS] run [PROGRAM FILE] [RUNTIME ARGUMENTS]
Description
The run command takes in the name of a single program file to execute and
runs it. The [RUNTIME ARGUMENTS] arguments may be omitted. These
arguments are passed to the script's main function as an array of string
arguments with the real-path of the script file always being the first.
12.1.3.The test Command🔗
Command Usage
tetral [OPTIONS] test [TEST DIRECTORY]
Description
The test command takes in a file path to a directory of Tetral program files and executes each after another.
Any errors thrown during the execution of the files are ignored and instead are checked against a list of "expected" errors.
Test Execution
For each valid test file:
- If the file is a Tetral source file, it is parsed and compiled. If any errors occurred during parsing or compilation, they are checked against a list of expected parsing errors.
- The test file's entrypoint is then executed as normal.
- Every function with a name prefixed with
test_is then executed. The function's name is considered to be the name of the test case. These functions do not need to be exported. - If any errors occurred, they are checked against a list of expected runtime errors. If all match, the test is considered passed. If there are more errors than expected or an expected error did not occur, or if an error did not match the expected one, the test is considered a failure and execution moves on to the next case.
- When all test cases in a single file have been executed, the
tests_shutdownfunction is called, if one exists. The shutdown function does not need to be exported. - If a test result directory is specified, test results, IR, ASTs and config information are dumped to the specified directory.
12.1.4.The help Command🔗
Prints text about the supported flags, commands and their arguments to the console. The version info is printed as well.
12.1.5.The version Command🔗
Prints information about the version of the Tetral language, compiler and interpreter in use.
12.2.Flags🔗
Certain flags will be referred to as "Value Flags," these are flags which require a value be provided after the flag's name. The value must be specified after an equals sign (U+003D.) For example:
--disable-warn=unused.variable
12.2.1.General Flags🔗
- The
--lib-directoryValue Flag The flag specifies a file path from which Tetral can load libraries during compilation or execution.
The flag can be specified multiple times to append more library directories.
- The
--no-implicit-stdFlag Specifies that the default
stdmodules should not be imported automatically.When specified, script files must explicitly import the
stdmodules.- The
--implicit-importValue Flag Specifies a module name that will be automatically imported without a source file having to explcitly state it's using symbols from that module.
12.2.2.Compiler Flags🔗
- The
--drop-line-numbersFlag - Instructs the compiler to omit all
PUSHLINEinstructions from the compiled output. - The
--error-warningsFlags - Instructs the compiler to treat all warnings as errors.
- The
--disable-warnValue Flag - Instructs the compiler to not report certain warnings
- The
--ignore-assertsFlag - Instructs the compiler to omit all
assertstatements from the compiled output. - The
--json-messagesFlag Instructs the compiler to print all messages in a JSON format instead of the default human-readable format.
This will result in messages with a schema defined in 13.2. Annex B. JSON Message Format
Regardless of this flag being used or not, the compiler will always print to either the standard output or the standard error output.
- The
--text-compile/-TFlag Writes a textual representation of the compiled Bytecode format to the output file instead of a binary format.
Textual IR files are not intended to be read and as such are not valid source or bytecode files. This flag is intended for debugging and human analysis of the compiled output.
- The
--output-file/-oValue Flag Tells the compiler the file path to output the compiled output to.
If this flag is not specified, then the file that is written to will be same path as the input file(s), except with the appropriate extension, depending on if the
--text-compileflag was specified, or not. (See 13.4. Annex D. Tetral File Extensions)
12.2.3.Test Command Flags🔗
- The
--verbose/-vFlag - Instructs the test executor to print more information to the standard output than it normally would, providing information about how long each test took to execute.
- The
--test-output-dir/-TOValue Flag Specifies a directory to which the test executor will dump information about each executed test to.
Each found test file will be executed and then information about the result will be printed to a dedicated directory nested inside the specified directory.
Printed information includes:
- A test file's AST, if parsed from a source file. (Dumped into a
ast.jsonfile.) - A test file's config. (Dumped into a
config.jsonfile) - A test file's compiled IR, if parsed from a source file. (Dumped into a
bytecode.tlirfile) - A summary of the test results. (Dumped into a
results.mdfile)
- A test file's AST, if parsed from a source file. (Dumped into a
13.Annexes🔗
13.1.Annex A. Class Implementation Details🔗
- Class methods are lifted out of classes and become global functions with the class' name and a dot character (U+002E) prefixed to the method's name.
- For non-
static methods, a hidden argument, thethis argument, is prepended to the method's arguments. - Class types exist in memory in the same way as struct types, with their functions becoming free-floating.
- If a class is exported, then all of its lifted functions are also exported.
- All constructors of a class are similarly transformed into global
functions with the type's name and a dot character (U+002E) prefixed and
the method name changed to
<constructor>. - A compiler-only statement is inserted into the body of each function, which calls the interpreter's allocator to allocate enough memory for the type, then the super constructor is invoked, if present, and then the rest of the constructor body.
13.2.Annex B. JSON Message Format🔗
Each logged JSON message is logged as a JSON object with the following properties.
level- A string determining the "level" of the message. Will always have one of
the following values:
FATALERRORWARNINFOTRACE
message_code- The message code, defined in 13.3. Annex C. Error and Warning IDs
message- Human readable, formatted message.
locationThis property may be omitted for certain messages, the presence of this property is not guaranteed.
If it is present, it will be a JSON object with the following properties:
start_lineandend_line- Line number range the message is referencing, starting at 1.
start_indexandend_index- Index range of the characters in the input the message is referencing, starting at 0.
13.3.Annex C. Error and Warning IDs🔗
This section details the IDs of compiler warning and errors. These IDs are
used to provide translations of the compiler and for the
--disable-warn argument.
| Message ID | Type | Default Value |
|---|---|---|
lexer.invalidcodepoint | Fatal Tokenizer Error | Invalid UTF-8 byte found in source file |
lexer.invalidocto | Fatal Tokenizer Error | Invalid octal sequence |
lexer.invalidhex | Fatal Tokenizer Error | Invalid hexadecimal sequence |
lexer.invalidbin | Fatal Tokenizer Error | Invalid binary sequence |
lexer.unclosedstring | Fatal Tokenizer Error | Unclosed string constant |
lexer.stringlinebreak | Fatal Tokenizer Error | Linebreak inside single-line string |
lexer.invalidescape | Fatal Tokenizer Error | Invalid escape sequence |
lexer.invalidunicode | Fatal Tokenizer Error | Invalid Unicode sequence |
lexer.invalidchar | Fatal Tokenizer Error | Character constant too long, may only contain one UTF-8 character |
parser.expected | Fatal Parser Error | Expected %0, found %1 |
parser.duplicateflag | Parser Error | Declaration modifier already set |
parser.funcdecl.trailingcomma | Parser Error | Illegal trailing comma in arguments declaration |
parser.funcdecl.nativebody | Parser Error | Native functions cannot have a function body |
parser.invalidlabeled | Fatal Parser Error | Illegal statement following loop label |
parser.typeexpr.invalid | Fatal Parser Error | Invalid type expression |
parser.module.nativenofrom | Parser Error | Native module declaration has no 'from' part to declare native library name. |
parser.module.invalidname | Fatal Parser Error | Expected module separator '::' to be followed by module name element. |
parser.call.trailingcomma | Parser Error | Illegal trailing comma in function call arguments. |
parser.call.malformed | Parser Error | Malformed call expression. |
parser.unexpected | Fatal Parser Error | Token not expected here, don't know how to parse. |
13.4.Annex D. Tetral File Extensions🔗
| File Extension | Description |
|---|---|
.tl or .tetral | Tetral source files |
.tlir | Tetral Bytecode file |
.tltir | Text representation of a Tetral Bytecode file |