A redesign of gluon's grammar and parsing types. Lives alongside v1
(github.com/accretional/gluon/pb, /lexkit, /metaparser) so we can
migrate callers incrementally; nothing in v2 depends on v1's proto
shapes (the CST implementation reuses v1's parser engine internally,
but that is an implementation detail invisible to v2 callers).
Go package root: github.com/accretional/gluon/v2
Generated proto types: github.com/accretional/gluon/v2/pb
Server + pure-Go entry points: github.com/accretional/gluon/v2/metaparser
v1 had a single LexDescriptor with a bag of named operator fields. v2
partitions a language's lex into two enums:
Delimiter— positional separators (WHITESPACE, DEFINITION, CONCATENATION, TERMINATION, ALTERNATION, …).Scoper— paired open/close brackets (OPTIONAL, REPETITION, GROUP, TERMINAL, COMMENT, …).
LexicalDelimiter { Delimiter kind; string symbol } and
LexicalScoper { Scoper kind; string begin; string end } carry the
concrete characters. SymbolDescriptor is a oneof over those two, and
LexDescriptor is just { name, repeated SymbolDescriptor symbols }.
v1 modelled a rule's RHS as a tree: ProductionExpression with nested
Sequence / Alternation / Optional / Repetition / Group wrapper
messages. v2 flattens this:
r = "a" , b | "c"
is encoded as
RuleDescriptor { name: "r", expressions: [
Production{ terminal: "a" },
Production{ delimiter: CONCATENATION },
Production{ nonterminal: "b" },
Production{ delimiter: ALTERNATION },
Production{ terminal: "c" },
]}
Delimiters sit between their siblings. Scopers wrap a nested
ScopedProduction { Scoper kind; repeated Production body }. The wire
form stays closer to the source and removes the
Sequence/Alternation/etc. wrapper layer.
DocumentDescriptor { name, uri, repeated TextDescriptor text } is the
top-level container. TextDescriptor is a oneof over four encodings:
AsciiChunk(repeated ASCII) — ASCII-only input (compact).UnicodeChunk(repeated int32) — reserved; never emitted by the current writers.unicode_string— UTF-8 string when the input contains non-ASCII.SourceLocation— by-reference chunk pointing at another document.
SourceLocation uses a uri (not inline text) deliberately, so the
message can never become a second home for source bytes.
Token's oneof over Delimiter delimiter / ScoperToken scoper has a
canonical "unset" form meaning literal text run. Because proto3 cannot
distinguish "oneof unset" from "oneof set to zero", five wire forms all
mean the same thing — producers emit form 1, consumers must accept all
five. This convention is called out with LOAD-BEARING banner comments
in tokens.proto and lex.proto; reordering the enum zero values or
reusing them as real roles would silently reclassify tokens as text.
Token carries an int32 offset outside the kind oneof (always
present) and an optional role. ScoperToken { oneof kind { Scoper lhs; Scoper rhs } } tags open/close sides. Using a oneof rather than a
{ kind, side } pair saves a tag+value on the wire per token.
Literal-run length is implied by the next token's offset, so the lexer
must emit a token (typically Delimiter.WHITESPACE) for every skipped
run — otherwise literal boundaries are lost.
Added during CST implementation. Tokens hold offsets, not text, so
matching a grammar terminal like "SELECT" against the source requires
the document itself. The RPC still does not re-lex — tokens remain an
input hint — but the source is the load-bearing input.
| File | Messages |
|---|---|
source_location.proto |
SourceLocation { uri, offset, length, line, column } |
text.proto |
TextDescriptor (oneof), AsciiChunk, UnicodeChunk |
document.proto |
DocumentDescriptor { name, uri, repeated TextDescriptor text } |
lex.proto |
Delimiter/Scoper enums, LexicalDelimiter, LexicalScoper, SymbolDescriptor, LexDescriptor |
tokens.proto |
TokenSequence, Token, ScoperToken |
grammar.proto |
Production (oneof), ScopedProduction, StringRange, RuleDescriptor, GrammarDescriptor |
ast.proto |
ASTDescriptor, ASTNode |
metaparser.proto |
service Metaparser, CstRequest, TransformRequest, TransformResponse, CompileRequest, CompileResponse |
astkit.proto |
service Transformer, Find/Replace/Filter Request+Response messages |
bytes → ReadBytes → TextDescriptor
string → ReadString → DocumentDescriptor
DocumentDescriptor (EBNF src) → EBNF → GrammarDescriptor
GrammarDescriptor + Document → CST → ASTDescriptor
ASTDescriptor (schema-shaped) → Compile → FileDescriptorProto
The lowering from ASTDescriptor to FileDescriptorProto lives in
v2/compiler and is exposed as the Compile RPC alongside the parse
pipeline. This is what v1 metaparser.Build used to do in one step;
v2 splits it into EBNF + CST + Compile so each stage is
independently addressable (and so non-EBNF front-ends can drop an AST
directly into Compile without re-lexing). compiler.GrammarToAST
is the bridge for callers that hold a GrammarDescriptor instead of
a schema-AST.
All six RPCs are implemented with pure-Go entry points, a unit test
file, and a _e2e_test.go that drives the gRPC stack via bufconn.
| RPC | Pure-Go entry | Unit tests | E2E tests |
|---|---|---|---|
ReadBytes |
ClassifyBytes(buf) (*TextDescriptor, error) |
16 | 13 |
ReadString |
WrapString(s) *DocumentDescriptor |
13 | 12 |
EBNF |
ParseEBNF(doc) (*GrammarDescriptor, error) |
13 | 9 |
CST |
ParseCST(req) (*ASTDescriptor, error) |
12 | 10 |
Transform |
Transform(ctx, *ASTDescriptor, script) (*TransformResponse, error) |
8 | 4 |
Compile |
compiler.Compile(ast, opts) (*FileDescriptorProto, error) |
17 | — |
go test ./v2/metaparser/ is green. The EBNF impl wraps v1's
lexkit.Parse and converts v1 ProductionExpression trees to v2's
flat Production list. The CST impl pretty-prints v2 rules back to
EBNF text, delegates to v1's lexkit.ParseAST, and converts the v1
AST to v2's shape. Both are drop-in replacements at the API level;
internally they ride on v1's parser engines.
The only external consumer today is proto-sqlite. It uses:
lexkit.LoadLex/lexkit.LoadGrammar— textproto loaders.lexkit.Parse(src, lex) → GrammarDescriptor— EBNF parse.lexkit.ToTextproto— serialize grammar to textproto.metaparser.Build(LanguageDescriptor) → FileDescriptorProto— the EBNF-grammar-to-compilable-proto compiler.
The rough path:
-
Add a v2 textproto loader. proto-sqlite keeps
sqlite-lex.textproto/sqlite-grammar.textproto; we need a v2-equivalent that parses into the v2 messages. Either translate the existing textprotos to v2 shape, or add aloadV2Textprotohelper inv2/metaparserthat unmarshals directly. -
Switch proto-sqlite's grammar pipeline to v2. In
lang/cmd/gengrammar/main.goandlang/cmd/genproto/main.go, replace the v1lexkit.Parsecall with a v2 client call toReadString+EBNF. The output textproto moves from v1GrammarDescriptorto v2GrammarDescriptor. -
Replace
metaparser.Buildwithcompiler.Compile. Done inv2/compiler:Compile(ast, opts) → *FileDescriptorProtowalks a schema-shapedASTDescriptor(kindsfile/rule/sequence/alternation/ …) and emits descriptor messages, deduplicating keyword terminals into empty messages.compiler.GrammarToASTbridges from aGrammarDescriptorfor callers (like proto-sqlite today) that hold a grammar rather than an AST. Exposed as theCompileRPC and as theprotoc://CompileTransform handler. -
Port v1 lexkit helpers callers still need. Only
LoadLex/LoadGrammar/ToTextprotoare used externally; the raw UTF8 helpers (Char,RuneOf,RawToToken,TokenToRaw) are v1-internal. Loaders need v2 equivalents; the UTF8 helpers can stay in v1 for now since v2 proto types don't use them. -
Delete the v1 package once callers are off it. v1
/pb,/lexkit,/metaparsercan be removed once proto-sqlite no longer imports them. Until step 3 lands, v2's CST implementation will continue to depend on v1 internally (as a parser engine), so v1 cannot be fully deleted yet — but it can be marked internal.
v2/astkit/ is the v2 replacement for the Go-ast-specific
github.com/accretional/gluon/astkit package. It offers plain-Go
helpers (Walk, Find, FindAll, ReplaceKind, Filter, Node,
Leaf, …) that operate on pb.ASTNode trees, plus an
astkit.Transformer interface whose method shape
(ctx, *Request) (*Response, error) is designed for codegen.OnboardDir.
Running OnboardDir("astkit", "v2/astkit") yields the
service Transformer proto — 6 rpcs (Find, FindAll, Count,
ReplaceKind, ReplaceValue, Filter). The checked-in v2/astkit.proto
is a lightly post-processed version of that output (fixed package
namespace to gluon.v2, added v2/ast.proto import, un-qualified
the ASTNode reference). The generated Go code lives in v2/pb/
alongside the other v2 protos.
The in-process implementation + gRPC adapter live in
v2/astkit/server/:
server.New() astkit.Transformer— pure-Go implementation.server.NewGRPCServer() pb.TransformerServer— gRPC adapter for wiring the service into agrpc.Server.
go test ./v2/astkit/... ./v2/metaparser/ is green (19 astkit unit +
11 server + 7 e2e tests).
Metaparser.Transform takes an ASTDescriptor and a textproto
ScriptDescriptor, runs the script through proto-expr's Protosh
runtime, and returns the final Data. The RPC wires v2/astkit's
Transformer methods in-process under the astkit://<Method> URI
scheme, so scripts can chain Filter / ReplaceKind / ReplaceValue /
Find passes on the tree without going through a separate gRPC hop.
Per-dispatch parameters ride in Data.type using a compact
k=v,k2=v2 convention (kind=whitespace, from=keyword,to=kw). The
initial ASTNode is pre-loaded into the ast register, so most
scripts open with request: { text: "ast" }. See
v2/metaparser/transform_test.go for examples.
The Compile RPC is also exposed inside Transform as the
protoc://Compile handler, so scripts can chain tree cleanups and
lowering in a single Transform call. Recognised params on the compile
step are package, go_package, file_name, and language; the
response Data carries the marshaled FileDescriptorProto with
type: "google.protobuf.FileDescriptorProto".
- Native v2 parser — replace the v1-parser shim inside
ParseCST. Would let us delete v1 wholesale. Not urgent; the shim is ~150 lines and works. TokenSequenceintegration — currentlyCstRequest.tokensis accepted but unused. Could drive pre-lex skipping of whitespace / comments once a native v2 parser exists.- Wire v2 into proto-sqlite — tracked as task #13 in this repo's task list.