To download the Operating System tOSh, you can do it from the Resources section of this same site.

Parts

This series of articles will have the following parts:

Armando Máquina de 1997Armando Máquina de 1997

Jumping to C

Sometimes, while writing this series, I ask myself whether I need to explain every concept behind every specialized term I use, like Kernel, and I never really know how deep I should go in my articles. Then I fill the terms with links and get over it.

As I’m writing this exact sentence, I’m asking myself: Do I have to explain what a Kernel is? I don’t think so. If you’re reading this, I assume you know a hell of a lot. And if you don’t, I’ll still fill every technical term I use with links so you can keep digging on your own.

Alright. We need to build a Kernel, because we’re already jumping into it, but we haven’t talked about anything related to how to start writing bytes for an operating system from C.

The first and perhaps most important thing you need to know about Kernel development is that you can’t just program like that, bam-bam-boom, grab a gcc or a clang, wham! slap a main.c in there, compile it, and shove the binary wherever the hell you feel like.

Noooooo! The first thing you need to do is set up a crosscompiler, that is, a compiler that lets you write binaries for other architectures without depending on the libraries that come with the target operating system or its ABI, because there is no ABI.

In other words, there are no libraries, no interfaces, no syscalls, nothing, because you don’t have a kernel, because you’re making one!

It might seem obvious to mention this, but if we stop for a second and think about it, we’re used to having hundreds of layers of abstraction that keep us from interacting with the CPU and solve hundreds of problems that get computed almost magically with one or two function calls. Writing a standard C printf is surprisingly complicated and is a __rabbit hole__ into which

not today, but someday, I’ll go down.

Having said all that, I strongly recommend reading the GCC Cross-compiler article on osdev.

KMain, Linkers and Bash

This is going to be our Kernel for now. The only thing it’s going to do is write the character "X" to the memory address VGA xb8000, which is the memory area where the video card is looking at its contents to replicate them on the corresponding monitor, and eventually we’ll end up crashing the CPU because we keep writing linearly all the way up to 2^32 - 1 or, even better, up to the limit of whatever memory we have installed in the hardware.

__attribute__((section(".text.kmain_entrypoint")))
int kmain(void) {

    char * vga = (char *) 0xb8000;
    *vga = 'X';

    while(1){
        *(vga++) ='X';
    };

    char * str = "end of kernel code.";
    return 1;
}

Compiling this code requires at least the following Makefile.

First, we define where we have the installation of our Cross-compiled version of GCC, which we compiled following the instructions in the GCC Cross-compiler article on osdev.

ASM=nasm
CROSS_DIR=~/opt/cross/bin
CC=$(CROSS_DIR)/i686-elf-gcc
AS=$(CROSS_DIR)/i686-elf-as 
LD=$(CROSS_DIR)/i686-elf-ld

Then we set the following compiler flags:

  • -m32 tells GCC to generate 32-bit code. That means, 32-bit general-purpose registers (eax, ebx, ebp, esp, edi, esi, etc…), 32-bit pointers, 32-bit ABI, instructions compatible with 32-bit protected mode.
  • -std=gnu99 is the C99 version of the C standard with GNU extensions, from 1999.
  • -Ttext 0x10000 is a linker flag that tells the linker that the .text section starts at address 0x10000.
  • -ffreestanding tells GCC not to assume that we have standard C runtime functions such as malloc(), printf(), etc…
  • -fno-pic disables Position Independent Code. Knowing that the entire Kernel starts at 0x10000 and onwards is enough.
  • -fno-asynchronous-unwind-tables and -fno-unwind-tables disable the metadata generated by GCC to reconstruct exceptions, tracebacks, stacktraces, in case of an error. This reduces the binary size, but also limits our debugging capabilities.
  • -march=i486 takes the instruction set way back to 486 to avoid advanced instructions (like, instructions from the Pentium family and up), limiting compatibility with older CPUs. For now, tOSh runs starting from i386.
  • -Map kernel.map is a linker option that lets us create a memory map and locate the offsets of each function and section in the binary.
CCFLAGS= -m32 -std=gnu99 -Ttext 0x10000 
CCFLAGS+= -ffreestanding -O0 -Wall -Wextra -fno-pic
CCFLAGS+= -fno-asynchronous-unwind-tables
CCFLAGS+= -fno-unwind-tables
CCFLAGS+= -march=i486 # avoid instructions like cmovna to be generated by gcc
LDFLAGS= -Map kernel.map

Finally, we define a standard Makefile. Nothing crazy here, we have .o, .map, and something special, which is our linker.ld file.

KERNEL_SRC=$(wildcard kernel/*.c)
KERNEL_OBJS=$(KERNEL_SRC:.c=.o)
KERNEL_BIN=kernel.bin

all: kernel 

clean:
	rm -rf *.bin 
	rm -rf *.o
	rm -rf *.map
	rm -rf *.img
	rm -rf kernel/*.o
	rm -rf *.lst

%.o: %.c
	$(CC) -o $@ -c $< $(CCFLAGS)


kernel: $(KERNEL_OBJS)
		$(LD) -o $(KERNEL_BIN) $^ $(LDFLAGS) -Tkernel/linker.ld

The Linker Script

When you compile, GCC generates some .o files that contain code and data, but each one is meant to be an independent chunk from the others. Also, each .o has several sections, such as .text, .rodata, .bss, etc…

What the linker does is take all those .o files generated when compiling one or more .c files and build a program by taking each of those OBJect Code files and positioning them, deciding exactly where each one ends up in memory.

Additionally, it resolves references. If the function pepito_delicioso() is programmed in mi_libreria_de_pepitos.c, that code will be in mi_libreria_de_pepitos.o, and if my kmain.c function calls it, the linker takes care of resolving the reference to that function.

The linker.ld file is a linker script. It’s a file that tells the linker not to place the OBJ files wherever the hell it feels like, but rather, you tell it where you want each thing placed in memory. By defining new sections or specifying where .text starts, you have more control over the layout of your program in memory.

OUTPUT_FORMAT("binary")
/*ENTRY(kmain)*/
SECTIONS
{
    /* the start address of the kernel's .text section */
    . = 0x10000;

    .text BLOCK(4K) : ALIGN(4K)
    {
        /* *(.text.prologue) */

        /* 
            set up the program entry point 
            kmain.c should have only one function,
            and it's going to be the first section of code executed
            when jumping to the kernel.
            in the generated kernel.map file, kernel/kmain.o should be 
            set at address 0x0000000000010000

            and void kmain(void); should be the first function written on top 
            of the program
        */
        *(.text.kmain_entrypoint)

        /* and later on, all the kernel objects */
        *(.text)
    }

    .rodata BLOCK(4K) : ALIGN(4K)
    {
        *(.rodata)
    }

    .data BLOCK(4K) : ALIGN(4K)
    {
        *(.data)
    }

    .bss BLOCK(4K) : ALIGN(4K)
    {
        *(.bss)
    }

    .idt BLOCK(4K) : ALIGN(4K)
    {
        *(.idt)
    }
    
    .dma BLOCK(4K) : ALIGN(4K)
    {
        *(.dma)
    }

    end = .;
}

Right at the very beginning, we define OUTPUT_FORMAT(\"binary\"). This makes the gcc linker generate a .bin directly instead of an .elf executable, with no headers, tables, debugging symbols, nothing. A flat binary.

Our Kernel knows nothing about binary formats or anything like that, and if we jumped into a .elf, the CPU would execute the bytes from the ELF format header and crash the CPU.

With . = 0x10000; inside SECTIONS, we tell the linker to start placing the binaries from that address. This points to .text, meaning that what we passed to the Makefile, where we tell it where the .text section starts, is redundant.

After that, each section has to be a block and also be aligned to "4-kilobyte pages".

This size, 4KB, is defined as a page in x86. Not only do the sections have to be 4KB blocks, they also have to be aligned to 4KB boundaries.

The first executable section is called .text.kmain_entrypoint, and what we do by defining it this way is that whichever C function has this attribute gets slapped right there, at the memory address 0x10000 that we already defined, making sure that function is the first one to execute.

If you look further up, in the C code I presented first, you’ll see the attribute above kmain() that the linker will take from the script so that it’s the very first thing to execute:

/*
0x10000 <-- arrancando en esta dirección de memoria
*(.text.kmain_entrypoint) <-- primero mete esta función
*(.text) <-- despues el resto de funciones
+-------------------------------+
| kmain_entrypoint              | <- PRIMERO
+-------------------------------+
| otras funciones .text         |
| configurar_cosas()            |
| contemplar_existencia()       |
| handlear_interrupciones()     |
| etc.                          |
+-------------------------------+
*/
__attribute__((section(".text.kmain_entrypoint")))
void kmain(void) { [...] }

The following sections, such as .bss, .rodata, and .data, are standard, but the ones that follow are not; they are specific to the design of my Kernel.

We’ll need the .idt section in the future when we define the interrupt table somehow like this:

    .idt BLOCK(4K) : ALIGN(4K)
    {
        *(.idt)
    }
__attribute__((section(".idt")))
struct idt_entry idt[256];

We do the same with the .dma section, which we’ll also use in the future to store some buffers:

.dma BLOCK(4K) : ALIGN(4K)
{
    *(.dma)
}
__attribute__((section(".dma")))
uint8_t dma_buffer[...];

So, roughly speaking, the Kernel’s memory layout would look something like this:

0x10000

+--------------------------+
| .text                    |
|                          |
| kmain()                  | <- primera función de todas, el entrypoint
| fs_init()                |
| fs_read()                |
| idt_init()               |
| ...                      |
+--------------------------+
| .rodata                  |
| strings / constantes     |
+--------------------------+
| .data                    |
| variables inicializadas  |
+--------------------------+
| .bss                     |
| variables sin init       |
+--------------------------+
| .idt                     |
| IDT                      |
+--------------------------+
| .dma                     |
| DMA buffers              |
+--------------------------+
|                          |
| end                      |
+--------------------------+

La Gotita

Finally, to join and glue all the pieces of the puzzle together with "La Gotita" and have an image to slap onto qemu, bochs, or a real machine, we need to generate an image containing the bootloader, stage1 and the kernel and each of their binaries positioned so that the bootloader and stage1 can copy things into memory, while making sure that during the process of recompiling and programming, the binaries don’t overwrite each other and are spaced far enough apart.

To achieve this, we had to sit down with a piece of paper and figure out how much enough space there should be, with "enough" being a variable that changed quite a bit throughout the development of this Kernel. More or less, this is the layout of the 3.5" 1.44MB Floppy I use to fit the Operating System onto (August 2026 update):

Sectors Offset Content Size
1 0x000000 Bootsector (boot.bin) 512 B
2–8 0x000200–0x000FFF Reserved space / Stage1 up to 7 sectors
9–136 0x001000–0x10FFF Kernel (kernel.bin) up to 128 sectors
137–2880 0x11000–0x167FFF Filesystem / remaining space ~1.37 MB
#!/bin/sh
echo "Cleaning up..."
make clean

echo "Compiling tOSh Kernel..."
make

kernel_size=$(stat -c %s kernel.bin)
kernel_sectors=$(( (kernel_size + 511) / 512 ))

MAX_KERNEL_SECTORS=128
KERNEL_FIRST_SECTOR=9
echo "kernel size: $kernel_size bytes, kernel sectors: $kernel_sectors"


if [ "$kernel_sectors" -gt "$MAX_KERNEL_SECTORS" ]; then
    echo "ERROR: kernel too large!"
    echo "Maximum: $((MAX_KERNEL_SECTORS * 512)) bytes"
    exit 1
fi

echo "Compiling bootsector and stage1..."
echo "taking into account the kernel size of $kernel_size bytes, which is $kernel_sectors sectors..."

echo "bootsector.asm -> boot.bin"
nasm -f bin bootsector.asm -o boot.bin


echo "Kernel first sector: $KERNEL_FIRST_SECTOR"

echo "stage1.asm -> stage1.bin with KERNEL_FIRST_SECTOR=$KERNEL_FIRST_SECTOR and KERNEL_SECTORS=$kernel_sectors"
nasm -D KERNEL_FIRST_SECTOR=$KERNEL_FIRST_SECTOR -D KERNEL_SECTORS=$kernel_sectors -f bin -l stage1.lst stage1.asm -o stage1.bin

stage1_size=$(stat -c%s stage1.bin)
stage1_sectors=$(( (stage1_size + 511) / 512 ))
echo "Stage1 size: $stage1_size bytes, $stage1_sectors sectors"

#nasm -f bin kernel/kernel.asm -o kernel.bin

echo "Creating floppy image..."
dd if=/dev/zero of=tOSh.img bs=512 count=2880

echo "Copying Bootsector at position 0, block 512 bytes..."
dd if=boot.bin of=tOSh.img conv=notrunc # copy the bootsector

echo "Copying stage1 at position 1, block 512 bytes..."
dd if=stage1.bin of=tOSh.img conv=notrunc bs=512 seek=1 # copy stage1

echo "Copying kernel at position $((KERNEL_FIRST_SECTOR -1)) (sector $KERNEL_FIRST_SECTOR), block 512 bytes..."
dd if=kernel.bin of=tOSh.img conv=notrunc bs=512 seek=$(($KERNEL_FIRST_SECTOR -1)) # copy kernel

Running this bash script ./build.sh creates the operating system on a floppy image ready to be booted on a machine or an emulator. So far, all we’re doing is writing X to memory (starting from the VGA mapping) until the CPU crashes from writing outside the memory limits.

This is the "basic" framework for being able to start developing a Kernel in C. From here on, we can keep configuring the CPU, program drivers for it, and take advantage of all the goodies of a high-level language like C, thus abandoning assembly.

References