Packing Binary Is Enjoyable, Really


Why would anybody do that?

I noticed somebody on twitter arguing that saving knowledge in JSON was apparently not what Real™ builders do.

Obviously, I needed to turn into a Real™ Developer too.

Turns out, the reply was binary.

Naturally, I had an excellent concept:

“How laborious may it’s to make my very own binary format?”

Surely it’s just a bit wb . ( ˶ˆᗜˆ˵ )

It was, sadly, not only a wb.

I ended up constructing a complete binary schema language that may shrink JSON payloads by 80%.

jBin

What is binary packing?

Say we obtained some knowledge

"hey globe"

then it could be translated in ascii to

104 101 108 108 111    32     119 111 114 108 100
 h   e   l   l   o   [SPACE]   w   o   r   l   d

So h turns into 104 in ASCII.

Since these ASCII values match inside 8 bits, every character takes up 1 byte.

h        e        l
01101000 01100101 01101100 

l        o        [SPACE]
01101100 01101111 00100000 

w        o        r
01110111 01101111 01110010 

l        d
01101100 01100100

so we are able to simply write it utilizing a lil bitta c:

FILE *f = fopen("file.bin", "wb");

unsigned char knowledge[] = "hey globe";
fwrite(knowledge, 1, sizeof(knowledge) - 1, f);

fclose(f);

Simple sufficient. Now let’s attempt writing 104.

Obviously, we may simply write 104 as ASCII characters:

'1' '0' '4' → 49 48 52

But that’s 3 bytes for a quantity that solely wants 1 byte.

So if we wish to save these 2 bytes, we’d like a way of telling the decoder, “hey, that is an integer, not a string.”

You may add a header, you’d solely be including an extra byte (properly relies on what number of varieties you bought.. hopefully you don’t have greater than 128 varieties…in case you do you bought greater points.

nice! lets simply use the primary byte to symbolize our sort and second to symbolize our knowledge.

[TYPE][DATA]

say 0 is int and 1 is string so 104 could be

00000000 01101000

and string could be..

00000001 01101000

oh wait…..

that might solely give is h

we’d like a option to symbolize totally different lengths of information. welp lets simply get one other byte. that ought to symbolize the size of our string.
so now our binary turns into

[TYPE][LENGTH][DATA]

nice! now we are able to symbolize our string like this:

 [TYPE]  [LENGTH] [DATA]
(STRING)  (11)
00000001 00001011 00...

h        e        l
01101000 01100101 01101100 

l        o        [SPACE]
01101100 01101111 00100000 

w        o        r
01110111 01101111 01110010 

l        d
01101100 01100100

GREAT! now we may pack each strings and ints collectively!
say we needed to symbolize "userid": 123

now you possibly can simply package deal all of it collectively

[TYPE:STRING][LENGTH:6][WORD:userid][TYPE:INT][LENGTH:-][DATA:123]

Great! we are able to symbolize 123 as a 1 byte quantity with 2 bytes of header.
however discover, we aren’t actually utilizing LENGTH subject for ints? why want it then? waste of bytes eh?

WELL… if we eliminate it, how does our binary reader know the place the header ends?

It wants some option to say “okay, the header is completed, begin studying the precise knowledge now.”

huh. what can we use to symbolize {that a} byte is ending.

A size byte for the header, maybe?

Ehh. That’s redundant. We’d be eradicating the size subject simply so as to add one other size subject.

But hey, we may use a bit within the header itself.

We may have one bit say:

I’m not the final byte within the header. There’s extra.

You would possibly assume: why not use the LSB?

Well, then we’d solely be capable of symbolize even numbers. Which is… not very best.

So we’ll use the MSB as an alternative.

so now our tag seems one thing like this:

[CONTINUATION BIT][7 BITS OF DATA]

If the continuation bit is 1, there’s one other header byte.
If it’s 0, the header is completed.

utilizing this, we are able to simply have our 123 be

[TYPE=0][DATA=123]

and if its a string.

[TYPE=1][LENGTH=11][DATA=104]...

so its of sort 1, size 11

however wait…what if the size is larger than 127? with 7 bits you possibly can solely symbolize as much as 127!

We use the identical factor! however for ints!

if the primary bit is 1 then the int continues.
128 might be written as:

10000001 00000000
^
MSB / continuation bit

(in large endian)

what the binary reader will do:

  • Reads the primary byte.
  • The MSB is 1, so there’s one other byte.
  • The remaining 7 bits are 1.
  • Reads the second byte.
  • Its MSB is 0, so that is the final byte.
  • Its remaining 7 bits are 0.
  • Combines the 2 7-bit values to get 128.

This is a type of varint (variable-length integer).

The encoding we’re utilizing right here is little-endian: the least-significant 7 bits come first.

10000000 00000001

Hey that is nice, innit? You can symbolize differing types in the identical binary and your binary parser will learn all of them accurately

But discover, We are storing this knowledge per subject.

[TYPE][DATA]
[TYPE][DATA]
[TYPE][DATA]

And most knowledge isn’t only a bunch of random values floating round. It’s normally structured.

Take a C struct:

struct {
    int i;
    char *s;
    int a[10];
}

this could be say on a 32 bit system.

[32-bit int] [32-bit pointer] [10 × 32-bit ints]

and we didn’t have so as to add headers everytime. as a result of we all know the kind of the info from the struct itself.

Hmm. I’m wondering if we are able to do that for our binary knowledge…

And sure, we are able to.

That’s what a schema is!

so for our struct our schema can simply be:

i: int
s: char *
a: listing(int)

The schema lets us know the sort with out storing the sort alongside each worth.

now our binary format doesn’t want to fret concerning the sort! it solely want to fret the dimensions of the info!

That’s what protobuf does

So lets take into consideration all of the totally different sizes of information we are able to have.

we obtained ints, we obtained floats, bools, strings.

We can deal with ints and bools as varints, whereas floats are fixed-width: f32 or f64.

Strings are totally different. We can’t simply encode their bytes as a varint, as a result of the bytes themselves are the precise knowledge we have to protect.

So as an alternative, we have to know what number of bytes belong to the string earlier than we begin studying it.

so now our encoding varieties are:

however wait! how can we entry our fields? like we cant simply go “gimme string” we might want to spacify which one. and no drawback lets simply symbolize every one with a quantity.
we are able to index them or have the consumer assign their very own numbers to handle them. that is what we’d like subject quantity for.

And discover one thing else: we solely have 4 potential varieties.

Four values match completely into 2 bits.

00 - varint 
01 - f32
10 - f64
11 - delimited

hey isn’t that neat. now what we may do is simply encode it…inside our subject quantity!

wait how?

some binary trickery..not likely.

You simply shove them collectively.

Move the sector quantity left by 2 bits to make room for our 2-bit sort, then OR the sort into these empty bits.

say your subject quantity is 10 and it’s a delimited sort.

00001010 (10)

left shift that by 2

00101000

OR it with our sort!

00101011

LOOK AT THAT! our 1 byte quantity tells us each what sort it’s and what its subject quantity is!

But what if we go FURTHER.

We’ve already packed the sort into the sector quantity.

Why cease there?

What if we may pack the size in too?

We can.

now our header can maintain

[FIELD NUMBER][LENGTH][TYPE]

ALL in a singular varint!!!! ◝(ᵔᗜᵔ)◜

Cool. Except for one factor

its simply a lot tedious work.
To pack a string, I’ve to inform it ‘that is delimited,’ give it the size, after which give it the bytes. Every. Single. Time.

PEASANTS DO THAT. Plebeians. we don’t try this.

So clearly, the answer is to jot down a complete schema language.
We declare our knowledge in a schema file, and let this system take care of all that tedious formatting and encoding nonsense.


Building the schema syntax

alright soo… we have to determine on a syntax that doesn’t suck your soul (taking a look at you protobuf)

subject numbers.. what are they? index proper. how do you index stuff in your grocery listing? you write quantity. merchandise
why not use the identical!

so one thing like this:

1. identify

we’d like the schema to symbolize the kinds. lets simply steal how different languages do it and do it like this:

1. identify: sort

neat huh.
however wait. how can we point out the top of a message(struct)

properly we may do {} however its not very good is it. why over complicate stuff its a listing. lists have an finish. lets have an finish.

message identify:
    1. identify:sort
    2. identify:sort
    3. identify:sort
finish

and no, indentations shouldn’t matter. its so annoying to work with languages the place indentation issues. its simply painful. lets simply not try this.

so..how will you establish the top of the road with out a semicolon? new line char?

I imply we may try this however what in the event that they needed to sort it in a single line, its ugly however say they wish to for no matter purpose. lets not limit that. however semicolons are ugly.

we may use the quantity!
if we see a quantity and a dot we simply take into account it a brand new entry!

neat huh.𐔌ˊᵕˋ𐦯

drawback…what if the sector is simply not discovered when the compiler is studying it?
we may crash…and we’d for all of the fields however we do need elective fields don’t we. lets simply go together with the apparent route and do one thing like this:

    quantity. identify: sort = defaultValue

and that’s elective. fairly intuitive.

now that we’re right here anyway lets take into consideration all of the totally different sorts of information individuals may symbolize…

properly we clearly obtained our entries of varieties.

oh we would want lists. we’d like some kind of option to decide between just a few issues so we would want an enum.
nice lets take into consideration these..

oh we may simply have them be perform like, that’s fairly intuitive.

    quantity. identify: construction(sort)

like:

    1. mates: listing(People)

oh wait however what’s individuals? oh it ought to be a message too. one thing like this:

message People:
    1. identify: string
    2. location: string
finish

huh.. we’d like customized varieties as properly… so within the language can we wish to have the whole lot be so as? ehh that’s fairly cringe later we may simply do a cross and put collectively all of our desk and simply have it seek advice from that struct.

hey due to this we are able to have self reference as properly. cuz once we pack it it’s going to simply be a pointer to the struct! so we are able to do one thing like:

message People:
    1. identify: string
    2. location: string
    3. mates: listing(People)

now for enums. cuz we’ve two passes we are able to simply put enums outdoors of message! it doesn’t matter the place it’s declared both! we are able to declare our enum one thing like this:

enum Name:
    quantity. identify 
    quantity. identify
    quantity. identify
finish

additionally protobuf does this factor the place it forces you to have the enum begin from 0
like can we REALLY want that? is the compiler so dumb it cant inform? lets simply have the compiler deal with that by default

oh wait…. individuals would possibly want some option to symbolize an enum however all the kinds usually are not the identical.
yup. that’s a union. lets simply add that no problemo

    quantity. identify: union(sort, sort, sort...)

oh wait.. our default worth. how would that work with unions.
if we’ve a union like this:

    quantity. identify: union(i32, i64) = 10

it’s ambiguous climate 10 is an i32 or an i64
so when subject is just not discovered what ought to we do?

we may repair this by having the default be subsequent to the sort.

    quantity. identify: union(i32 = 10, i64)

now if a the sector is just not discovered will probably be an i32 with worth 10

okay properly individuals would wanna map stuff as properly…so we add

    quantity. identify: map(key_type, value_type)

lets simply add syntax for declaring packages and importing stuff:

package deal "package deal identify"
import "package deal"

hey would you take a look at that we’ve some very neat syntax. lets simply put all of it collectively:

package deal "com.recreation.core"
import "math.jbin"

enum Activity:
    1. energetic
    2. inactive
finish

message Player:
    1. identify: string
    2. well being: i32 = 100
    3. weapons: listing(string)
    4. connections: listing(Player)
    5. activeStatus: Activity
    6. stock: map(string, i32)
    7. stability: union(string="empty", i32)
finish

oh wait…we’ve an entire language now….

anyway. that is the equal of it in .proto

syntax = "proto2";

package deal com.recreation.core;

import "math.proto";

enum Activity {
  ACTIVE = 1;
  INACTIVE = 2;
}

message Player {
  elective string identify = 1;
  elective int32 well being = 2 [default = 100];
  
  repeated string weapons = 3;
  repeated Player connections = 4;
  
  elective Activity active_status = 5;
  map stock = 6;
  
  oneof stability {
    string empty = 7;
    int32 quantity = 8;
  }
}

take a look at that. ew.


Time to truly implement this

To implement this I naturally selected C++.

By “naturally,” I imply I needed one thing I may placed on a resume and wasn’t within the temper to battle the Rust borrow checker. I’ve aged sufficient.

Okay. Enough designing. Time to truly make the factor work.

First, we have to encode our binary.

Encoding and Decoding our Data

To encode a varint, it’s mainly simply this:

void encodeVariant(std::vector<uint8_t> &buffer, uint64_t worth) {
    whereas (worth >= 128) = 128;
        buffer.push_back(decrease);
        worth >>= 7;
    
    uint8_t decrease = worth & 127;
    buffer.push_back(decrease);
}

This builds our varint. If the worth is >= 128, we take the bottom 7 bits, set the MSB to 1 to say “there’s extra,” and append it to the buffer.

Then we shift the worth by 7 bits and repeat.

For the ultimate byte, we go away the MSB at 0.

fairly easy. and we are able to decode it utilizing this:

uint64_t decodeVariant(const std::vector<uint8_t> &buffer, size_t &offset) = forged;
    return consequence;

Now that we are able to encode and decode our varints utilizing LEB128 lets encode and decode some tags(our header metadata we mentioned about)

void encodeTag(std::vector<uint8_t> &buffer, uint32_t subjectNumber,
               wiretype wiretype) = static_cast<uint64_t>(wiretype);
    encodeVariant(buffer, val);


void decodeTag(const std::vector<uint8_t> &buffer, size_t &offset,
               uint32_t &outFieldNumber, wiretype &outType) {
    uint64_t val = decodeVariant(buffer, offset);
    outType = static_cast<wiretype>(val & 3);
    outFieldNumber = val >> 2;
}

And we are able to add a few helpers for strings:

void encodeString(std::vector<uint8_t> &buffer, uint32_t subjectNumber,
                  const std::string &textual content) {
    encodeTag(buffer, subjectNumber, wiretype::Delimited);
    encodeVariant(buffer, textual content.dimension());
    buffer.insert(buffer.finish(), textual content.start(), textual content.finish());
}

std::string decodeString(const std::vector<uint8_t> &buffer, size_t &offset) {
    uint64_t dimension = decodeVariant(buffer, offset);
    std::string consequence =
        std::string(buffer.start() + offset, buffer.start() + offset + dimension);
    offset += dimension;
    return consequence;
}

Encoding a string is now simply three issues: write its tag, write its size, then write its bytes.

consider it or not, that’s mainly all our core engine performed! the remaining is simply the compiler and the json conversion stuff! ٩(^ᗜ^ )و ´-

Great! now that we’ve written our binary encoding… lets make a compiler ought to be easy…proper?

Building Compiler

Yes it’s a compiler. cease it I don’t wanna name it a transpiler it’s a compiler. the definition is:

compiler is a pc program that interprets supply code written in a single programming language (the supply language) into one other programming language (the goal language), whereas preserving the precise that means and conduct of the unique code.

ours does that. gtfo compiler individuals. my program takes my schema and interprets it into both a binary or code.

The drawback

Our language is fairly cute. Pretty slick.

Unfortunately, the pc has no concept what any of it means.

It’s simply bytes.

so lets assign that means to those symbols!

lets take our syntax:

    1. identify:string

so its within the construction of:

quantity → dot → identifier → colon → sort

after which we are able to use this to construct our illustration of those entries. that’s what a Lexer does

Building the Lexer

Yes lexer is like what you assume it’s. its only a large loop with a bunch of if statements.
All it does is learn our file and spit out these tokens.(not the ai form)

enum class TokenKind {
    Keyword_Message,
    Keyword_Enum,
    Keyword_Optional,
    Keyword_Map,
    Keyword_Union,
    Keyword_End,
    Keyword_Package,
    Keyword_Import,
    Identifier,
    Number,
    StringLiteral,
    Comment,
    Equals,
    Colon,
    Comma,
    Dot,
    LParen,
    RParen,
    EndOfFile
};

so our file

message User:
    1. identify: string
    2. id: i32

the tokenized output could be:

    [Keyword_Message]
    [Identifier: "User"]
    [Colon]
    [Number: 1]
    [Dot]
    [Identifier: "name"]
    [Colon]
    [Identifier: "string"]
    [Number: 2]
    [Dot]
    [Identifier: "id"]
    [Colon]
    [Identifier: "i32"]
    [EndOfFile]

Pretty neat. Now the parser doesn’t must decipher a bunch of uncooked characters. It can simply work with these tokens.
parser seems at these after which really emits the AST (summary syntax tree).

What is AST?

Its simply how we symbolize our program knowledge in graphLang it was a literal tree node that held all the info and all knowledge was only a single form.
For this one we are able to outline our schema as the foundation. we simply have two sorts of messages rn messages and enums so our head might be outlined as:

struct Schema {
    std::vector<EnumDef> enums;
    std::vector<MessageDef> messages;
};

so our Schema would be the root node and the tree will seem like this:

    (Schema)
    /      
(enums) (messages)

Building the AST

soo…. what’s enums and messages? from the definition above you possibly can see its a vector of structs.
our messages might be outlined as:

struct MessageDef {
    std::string identify;
    std::vector<Field> fields;
    int line = 0;
    std::string remark = "";
};

after which our Field as:

struct Field {
    uint32_t quantity;
    std::string identify;
    DataKind sort;
    bool isOptional = false;
    std::string defaultValue = "";
    int line = 0;
    std::string remark = "";
};

now our AST seems like this:

Schema
└── messages
    └── MessageDef
        ├── identify
        ├── line
        ├── remark
        └── fields
            └── Field
                ├── quantity
                ├── identify
                ├── sort
                ├── isOptional
                ├── defaultValue
                ├── line
                └── remark

and all that’s left is our enum definition:

struct EnumDef {
    std::string identify;
    std::vector<EnumEntry> entries;
    int line = 0;
    std::string remark = "";
};

and our enum entry:

struct EnumEntry {
    uint32_t quantity;
    std::string identify;
    int line = 0;
    std::string remark = "";
};

Great! we obtained all our items. our AST seems like this now:

Schema
├── Enums
│   └── EnumDef
│       ├── identify
│       └── entries
│           └── EnumEntry
│               ├── quantity
│               └── identify
│
└── Messages
    └── MessageDef
        ├── identify
        └── fields
            └── Field
                ├── quantity
                ├── identify
                ├── sort
                ├── isOptional
                └── defaultValue

Building the parser

we use this and construct our parser. and sure. our parser is only a loop with if statements (properly technically recursive first rate or no matter however recursion is only a type of iteration).

it’s really fairly easy the entire loop:

Schema Parser::parse() {
    Schema s;
    whereas (!isAtEnd()) {
        if (peek().sort == TokenKind::Comment) {
            devour();
            proceed;
        }
        if (peek().sort == TokenKind::Keyword_Package) {
            devour();
            if (peek().sort == TokenKind::StringLiteral) {
                s.package dealName = devour().worth;
            } else {
                error("Expected string literal after package deal");
            }
        } else if (peek().sort == TokenKind::Keyword_Import) {
            devour();
            if (peek().sort == TokenKind::StringLiteral) {
                s.imports.push_back(devour().worth);
            } else {
                error("Expected string literal after import");
            }
        } else if (peek().sort == TokenKind::Keyword_Message)
            s.messages.push_back(parseMessage());
        else if (peek().sort == TokenKind::Keyword_Enum)
            s.enums.push_back(parseEnum());
        else
            error("Unexpected token in international scope");
    }
    return s;
}

and for every of the kinds its only a bunch if circumstances checking after which returning the AST node.
devour() returns the present token and strikes the pointer over by one and peek() provides you the following token knowledge with out transferring the pointer

GREAT! lets try it out.

    message User:
        1. identify: adfasdfasdfasdfasd
        2. id: 012349

lets see what our parser says:

“Mighty good mate! appears wonderful innit? desire a cuppa?”

yeah..so we’d like a factor that checks the consumer isn’t simply syntactically appropriate but additionally the shit is smart.

so we’d like a reality checker, a twitter group notice if you’ll.

And that’s what our semantic analyzer does!

Building the Semantic Analyzer

So to validate stuff we shall be doing it in two passes.

  • First cross we construct all of our symbols and put them right into a desk. (Basically simply all of the legitimate stuff which might be identifiable)
  • Second cross we use that the desk we constructed validate the AST.
void SemanticAnalyzer::analyze(const Schema &schema) {
    buildSymbolTable(schema);
    validateEnums(schema);
    validateMessages(schema);
    validateCyclicDependencies(schema);
}

Because we do that in two passes, declaration order doesn’t matter. Which is fairly neat.

So throughout validateMessages , when it seems at
adfasdfasdfasdfasd , it checks our image desk,
realizes that sort doesn’t exist, and throws an error . It additionally checks
that you just didn’t do one thing silly like use subject
quantity 1 twice, or use a listing as a map key.

Great! Look at what we obtained up to now!

  • A Lexer that chops the whole lot up into tokens
  • A Parser that takes the tokens and builds the AST
  • A Semantic Analyzer that validates the AST and throws errors
  • An encoder/decoder utilizing LEB128.

Now all that’s left to do is to do two issues.

  • Convert JSON recordsdata into binary dynamically
  • Codegen from the schema to allow them to use them of their language natively

Dynamic Packer

Now for the dynamic packer. It takes a schema, takes a JSON file, and builds our binary.

For JSON, I simply used nlohmann/json. Parsing JSON on prime of the whole lot else would’ve been a very pointless aspect quest.

So we’ve our JSON loaded in reminiscence. We have our
AST loaded in reminiscence. Now we simply stroll them collectively.

When the JSON parser sees the
key “id”: 123 , it doesn’t know what to do. But it
asks the AST! The AST says,

“Oh, id ? That’s subject quantity 2, and it’s an i32 .”

So the Packer simply calls our encodeTag() and encodeVariant()
capabilities and spits the bytes into our buffer.

Notice what we didn’t do? We
didn’t generate any C++ code. We didn’t compile any
wrappers. We simply learn the JSON, learn the schema, and
constructed the binary dynamically.

take that protobuf

Lets try it out!

{
    "identify": "Pranav",
    "well being": 100
}

If we save this as JSON, with areas and quotes,
it’s about 35 bytes

lets pack it in binary.

and that’s…

NINE BYTES

hell yeahh take a look at that!!!

74.29% REDUCTION

it could be even increased if we didn’t have strings and stuff.
6 of the 9 bytes are used for the phrase “Pranav”
however apart from the purpose.

we’ve diminished our dimension of storage by a LOT.

And decoding is means easier too. The binary already tells us what every subject is meant to be, so the decoder doesn’t must take care of JSON’s syntax and kind illustration.

But what in case you don’t wish to pay the price of dynamically wanting the whole lot up at runtime? What in case you simply need regular structs in your language?

That’s the place AOT code technology is available in.

Codegen

Hmm… how would one generate code? I imply we’ve our AST so we all know what it seems like and what it’s semantically so like its simply translating that into our language..

you possibly can write a cCodeGen() func and add it to our class and name.

if we needed a generator for python? oh that’s straightforward! i’ll simply add a pythonCodeGen!

oh I want a jsCodeGen() okay look. too far. why are you writing javascript. however regardless.

But apparently the shopper is all the time proper of their language preferences or no matter.
okay now this class is getting a wee bit too bloated for my liking.

That’s precisely why we’d like a customer sample!

Visitor Pattern

So what’s a customer? in hindsight, design sample cope for ones who’s language doesn not have algebraic datatypes. So why did I take advantage of it regardless that cpp has the std::variant? Idk I learn it in an article as soon as. I needed ot study it.

Anyway. so a customer sample is only a lil handshake the caller and callee do. we get sort security from it. in our implementation our caller passes itself in and turns into the customer. the callee or the acceptor accepts the customer and does a little bit func name on the customer passing itself in triggering a double dispatch. fairly neat..however its simply a lot psychological overhead for a easy drawback.

to perform this we might want to edit our structs within the ADT to additionally maintain a perform.

void settle for(SchemaVisitor &customer) const;

and we create an summary class with a bunch of digital strategies that our mills implement:

#pragma as soon as
#embrace "Schema.hpp"

class SchemaVisitor {
  public:
    digital ~SchemaVisitor() = default;
    digital void go to(const Schema &s) = 0;
    digital void go to(const MessageDef &m) = 0;
    digital void go to(const EnumDef &e) = 0;
    digital void go to(const EnumEntry &ee) = 0;
    digital void go to(const Field &f) = 0;
};

so our c generator for instance seems like this:

#pragma as soon as
#embrace "SchemaVisitor.hpp"
#embrace 
#embrace 

class CGenerator : public SchemaVisitor {
  non-public:
    std::ostream &out;
    std::string currentEnumName;
    const Schema* presentSchema = nullptr;

  public:
    CGenerator(std::ostream &outputStream) : out(outputStream) {}

    void go to(const Schema &schema) override;
    void go to(const MessageDef &message) override;
    void go to(const EnumDef &enumDef) override;
    void go to(const EnumEntry &ee) override;
    void go to(const Field &subject) override;
};

so our C generator simply overrides these funcs and in these visits it does this:

void CGenerator::go to(const EnumDef &enumDef) {
    currentEnumName = enumDef.identify;
    out << "typedef enum {n";
    for (const auto &ee : enumDef.entries) {
        ee.settle for(*this);
    }
    out << "} " << enumDef.identify << ";nn";
}

once we do ee.settle for(*this) it does this:

void EnumEntry::settle for(SchemaVisitor &customer) const { customer.go to(*this); }

so it simply calls CGenerator.go to(ee) so the go to for surroundings entries known as in CGenerator.

void CGenerator::go to(const EnumEntry &ee) {
    out << "    " << currentEnumName << "_" << ee.identify << " = " << ee.quantity << ",n";
}

and that is performed for every sort. and that’s how the code is generated.

GREAT! lets try it out..

oh wait. we cant. we don’t have a means to try this. we’d like a cli.

Building CLI

For the CLI I used jarro2783/cxxopts cuz I didn’t wish to do handbook parsing. will probably be a enjoyable challenge each this and the json I’ll do them someday else however for now I used these.

And constructing the cli was fairly straightforward from this library you get a bunch of stuff that simply works.

I simply wrote my choices.

choices.add_options()("command", "Command to run (e.g. construct, pack)",
                              cxxopts::worth<std::string>())(
            "enter", "Input schema file", cxxopts::worth<std::string>())(
            "o,out",
            "Output (goal language for construct, or output binary file for "
            "pack)",
            cxxopts::worth<std::string>())("j,json",
                                           "Input JSON file (for pack command)",
                                           cxxopts::worth<std::string>())(
            "m,msg", "Root message identify to pack (for pack command)",
            cxxopts::worth<std::string>())("h,assist", "Print utilization");

        choices.parse_positional({"command", "enter"});
        auto consequence = choices.parse(argc, argv);

and it really works. it was nice. and these choices have been dealt with in a bunch of if statements(now that I give it some thought most of this challenge has simply been a loop and a bunch of if statements)

if (command == "construct") {
            if (!consequence.rely("enter")) {
                std::cerr << "Error: No enter file specified." << std::endl;
                return 1;
            }
            if (!consequence.rely("out")) {
                std::cerr << "Error: --out flag is required (e.g., --out c)."
                          << std::endl;
                return 1;
            }

            std::string inputFile = consequence["input"].as<std::string>();
            std::string targetLang = consequence["out"].as<std::string>();

            std::ifstream file(inputFile);
            if (!file.is_open()) {
                std::cerr << "Error: Could not open file " << inputFile
                          << std::endl;
                return 1;
            }

            std::stringstream buffer;
            buffer << file.rdbuf();
            std::string schemaText = buffer.str();

            std::vector<Token> tokens = tokenize(schemaText);
            Parser parser(tokens);
            Schema schema = parser.parse();

            SemanticAnalyzer analyzer;
            analyzer.analyze(schema);

            if (targetLang == "c") {
                CGenerator cGen(std::cout);
                schema.settle for(cGen);
            } else if (targetLang == "py" || targetLang == "python") {
                PythonGenerator pyGen(std::cout);
                schema.settle for(pyGen);
            } else {
                std::cerr << "Code technology for '" << targetLang
                          << "' is just not supported but!" << std::endl;
            }

anyhow. lets try it out.

we are going to write our schema as this:

package deal "com.mmo.recreation"

enum Faction:
    1. alliance
    2. horde
    3. impartial
finish

message Vector3:
    1. x: i32
    2. y: i32
    3. z: i32
finish

message StockItem:
    1. itemId: i32
    2. amount: i32
    3. isSoulbound: bool
finish

message Character:
    1. id: i64
    2. identify: string
    3. stage: i32
    4. faction: Faction
    5. place: Vector3
    6. stock: listing(StockItem)
    7. attributes: map(string, i32)
finish

and run construct with output as c… and..

#pragma as soon as
#embrace 
#embrace 
#embrace 
#embrace 

typedef struct Vector3 Vector3;
typedef struct StockItem StockItem;
typedef struct Character Character;

typedef struct {
    StockItem* knowledge;
    size_t size;
    size_t capability;
} jbin_list_InventoryItem;

typedef struct {
    char** keys;
    int32_t* values;
    size_t size;
    size_t capability;
} jbin_map_char_ptr_int32;

typedef enum {
    Faction_alliance = 1,
    Faction_horde = 2,
    Faction_neutral = 3,
} Faction;

struct Vector3 {
    int32_t x;
    int32_t y;
    int32_t z;
};

struct StockItem {
    int32_t itemId;
    int32_t amount;
    bool isSoulbound;
};

struct Character {
    int64_t id;
    char* identify;
    int32_t stage;
    Faction faction;
    Vector3 place;
    jbin_list_InventoryItem stock;
    jbin_map_char_ptr_int32 attributes;
};

LOOK AT THAT!! zero dependency left from this system. we offer all of the stuff it wants from our schema!!!

(notice: that isn’t the complete c file that was generated. there have been additionally a number of pack/unpack capabilities for particular person messages, setters and getters and encode/decode funcs)

we are able to additionally encode to binary by passing in a json file.

{
    "id": 123456789,
    "identify": "LeroyJenkins",
    "stage": 60,
    "faction": "alliance",
    "place": {
        "x": 100,
        "y": 200,
        "z": 300
    },
    "stock": [
        { "itemId": 999, "quantity": 1, "isSoulbound": true },
        { "itemId": 45, "quantity": 100, "isSoulbound": false }
    ],
    "attributes": {
        "energy": 120,
        "agility": 45
    }
}

and the dimensions of this json is 401 bytes

lets pack it into binary. and the file dimension is..

80 bytes!

EIGHTY PERCENT REDUCTION

80.05%

Conclusion

So yeah. I began out simply wanting to avoid wasting
JSON out of spite, and I by chance constructed a lexer,
a parser, an AST, a semantic analyzer, a dynamic
binary packer, and a multi-language code generator.
And truthfully? Packing binary is enjoyable, really.

anyhow, checkout the repo.

jBin



Source link