gatewai-dev/gitframes: Code-first video, rendered natively on WebGPU. · GitHub


Compositing, movement graphics and 3D for code-first video — one npm bundle that AI brokers drive with code.

npm
license
status
discord
youtube
node
engine
gpu
vision

⚠️ Beta: gitframes is below energetic improvement. APIs might change between releases and a few options could also be incomplete or unstable.

Gitframes is constructed for coding brokers. It packs the work folks often cut up throughout three desktop apps (Photoshop-inspired compositing and VFX, After Effects-style movement, typography and keyframing, and Blender-style 3D scenes, cameras and fashions) into one light-weight npm bundle. Your agent writes a TypeScript composition, checks frames, and renders an MP4, and no person has to put in or license a multi-gigabyte inventive suite.

Code-first video as pure software program engineering — no headless browser, no DOM reflow, no screenshot pipeline.
Renders instantly on GPU {hardware} through Dawn / WebGPU / Metal / Vulkan in Node.js and trendy WebGPU browsers.

Every body of those movies is rendered by gitframes from TypeScript in examples/. Click a nonetheless to observe it on YouTube.

Note

Using an AI coding agent? Install the gitframes expertise in a single line.

Claude Code

/plugin set up gitframes

Any different agent (Codex, Cursor, Hermes, Gemini CLI, Copilot, and extra)

npx expertise add gatewai-dev/gitframes

See Agent Skills & Plugins for particulars.



Modern automated video era is often constrained by the architectures of general-purpose internet browsers: course of overhead, non-deterministic DOM format reflows, and gradual screenshot seize. Gitframes treats video composition as software program engineering:

Pillar What it means
🚀 Zero Headless-Browser Overhead No Puppeteer, no Chromium IPC, no web page.screenshot(). Gitframes talks straight to native GPU gadgets through Dawn/WebGPU and hardware-encodes with @napi-rs/webcodecs.
🎯 Deterministic Frame-Accurate Clock Absolute body clocks, discrete pattern factors, and frame-accurate audio BeatGrids. No floating timers, no drift, no dropped frames.
🔠 Analytic, Resolution-Independent Type The Slug algorithm evaluates glyph contours per-pixel in WGSL — no texture atlases, no scaling artifacts, razor-sharp from 10 px to 10,000 px.
🎨 Photoshop-Inspired Tonal & Spatial VFX 50+ modular GPU shaders: Curves, Levels, Selective Color, 3D LUTs, Halftone, Film Grain, Unsharp Mask, Mesh Warp, and Screen-Space Relighting.
🧊 Unified 3D & 2D Depth Compositing Nest 2D flex/field bushes inside 3D homography planes, multiplane rigs, and meshes (OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF), with PBR glass and SSAO.
🔊 Built-in Procedural Audio DSP Multi-track soundtracks, deterministic procedural transition SFX (whoosh, affect, riser), and reactive indicators that drive visuals from audio.
👁️ On-Device Neural Vision Object monitoring, occasion segmentation, multi-person pose, and particular person mattes from Apache-2.0 ONNX fashions — feeding reactive indicators and not using a spherical journey to disk.
☁️ Cloud-Native & CI/CD Ready ~200–400 MB RAM per employee (vs. 2–4 GB for Chromium), ultimate for serverless GPU render clusters (AWS G4/G5, Modal, RunPod, Kubernetes).


Architectural Comparison: Gitframes vs. Remotion vs. Hyperframes

Developers producing video programmatically generally weigh Remotion (React/Chromium) or Hyperframes (Canvas2D/SVG internet animation). The matrix under compares the basic engineering dimensions.

Detailed Comparison Matrix

Capability / Dimension Gitframes Remotion Hyperframes
Underlying Engine Native WebGPU (WGSL compute & render pipelines through Dawn / Metal / Vulkan) Chromium / Puppeteer (React DOM, HTML/CSS format) Canvas2D / WebGL / SVG (browser or Node Skia)
Rendering Architecture Direct {hardware} framebuffer rendering & {hardware} video encoding (@napi-rs/webcodecs) Spawns headless Chrome; captures frames through CDP / web page.screenshot() Software or {hardware} 2D canvas context
Throughput 60–120+ FPS (real-time to faster-than-real-time GPU execution) 5–20 FPS (DOM reflow, IPC, rasterization) 20–40 FPS (CPU draw instructions / JS)
Memory Footprint ~200–400 MB per render (zero browser) 1.5–4.0 GB+ per employee (Chromium + V8 DOM heap) ~500 MB–1 GB (Skia/Canvas bindings)
Typography Engine Slug GPU — analytic Bézier analysis in WGSL, infinite zoom, After Effects selectors Browser DOM textual content (CSS fonts, rasterized, blurry below 3D transforms) Canvas2D / path textual content (CPU-rasterized glyphs)
2D VFX & Post-Processing 50+ WebGPU shaders (Curves, Levels, Selective Color, 3D LUT, Film Grain, Halftone, Liquify, PBR Glass, Relight) CSS Filters or customized WebGL canvas wrappers Basic Canvas2D composites and 2D filters
3D Graphics & Depth Native 3D scene graph — LookAt/Turntable digicam, multiplane, skinning (OBJ/FBX/glTF), SSAO, PCSS, DoF None built-in (embed Three.js/Fiber inside React DOM) Minimal 2.5D layers; no unified mesh pipeline
Motion Blur & Physics Physical 180° shutter velocity buffers in MRT + closed-form spring kinematics CSS transitions / JS interpolation; artificial blur hacks Frame interpolation or handbook multipass
Audio Engine & DSP Native audio DSP & procedural SFX (multi-track mixing, beat grids, reactive indicators) playback; fundamental quantity curves Basic static audio playback
Charts & Data Viz Layer.chart — line, space, bar, scatter, candlestick, pie and donut charts constructed from native vector nodes, with staggered reveal animations DOM chart libraries (Recharts, Chart.js) Custom canvas draw operations
AI & Computer Vision On-device ONNX imaginative and prescient — COCO-80 detection + occasion masks (RTMDet-Ins), COCO-17 pose (RTMO), particular person mattes (Selfie Segmenter); WebGPU tensor conditioning (Canny, depth-to-normals, optical stream, deflicker) External pre-rendered property; no native GPU tensor conditioning External pre-rendered property
Headless Verification FrameGrid contact sheets, single-frame snapshots, Skia MSE pixel-invariant assertions Playwright/Puppeteer visible snapshots Manual body inspection / canvas diffing
Docker / Cloud Portability Compact (~500 MB slim picture with native GPU/Vulkan drivers) Heavy (~2–3 GB with Chromium, fonts, X11/Mesa) Moderate container dimension


Key Features & Capabilities

1. Slug GPU Vector Typography & After Effects Animators

Traditional textual content depends on CPU rasterization or low-res SDF atlases that soften below 3D digicam sweeps. Gitframes integrates the Slug algorithm (SlugPipeline):

  • Analytic GPU analysis — WGSL fragment shaders resolve precise cubic/quadratic Béziers per-pixel. Glyphs keep sharp at 10 px or 10,000 px with zero CPU re-rasterization.
  • After Effects–fashion animators — vary selectors (sq., ramp_up, ramp_down, triangle, clean), easeHigh/easeLow curves, and seeded PRNG character shuffling (TextAnimator).
  • Human typing cadence — weighted punctuation delays (commas 3×, sentence ends 5.5×, newlines 7×) and trailing scramble decision (TypewriterAnimator).
  • 3D volumetric formations — map textual content onto cylindrical drums, logarithmic vortex spirals, and double-helix ribbons with surface-normal banking (evaluateVolumetricFormation).
  • Dynamic main & skew — area-preserving unimodular shear and accordion line-leading anchored to baseline, heart, or high.

2. Photoshop-Inspired WebGPU 2D VFX (50+ Shaders)

A complete suite {of professional} picture/video shader nodes in nodes/ and packages/webgpu-renderers:

  • Tonal grading — Curves (RGB/R/G/B spline), Levels (black/white level, gamma, output), Shadows/Highlights, Selective Color (CMYK gamut isolation), 3D Cube LUT (ApplyLUT).
  • Stylization & grain — Film Grain (Gaussian emulsion with spatial seed variation), Halftone (mono/RGB/CMYK, adjustable dot form & angle), Gradient Map, High Pass.
  • Optics & lens — Bilateral Gaussian Blur, Unsharp Mask, Vignette, Refraction Caustics, PBR Glassmorphism with chromatic dispersion (PBRGlass).
  • Distortion & warping — Displacement Maps, Liquify, Mesh Warp, Corner Pin homography.

3. Unified 3D Scene Graph, Camera & Mesh Shading

  • Calibrated digicam rig — LookAt and Turntable cameras (Camera3D) calibrated so z = 0 matches 2D canvas pixel coordinates 1:1.
  • 3D format primitives — Layer3D.cube, carousel, prism, airplane, grid with unified depth-buffer testing.
  • Zero-dependency mannequin parsers — OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF.
  • Skeletal animation & shading — 128-bone Linear Blend Skinning, Blinn-Phong & PBR multi-light shading, PCSS/Poisson contact shadows, SSAO, and optical DoF.
  • Physical movement blur — 180° shutter movement blur with per-vertex velocity vectors packed into rg16float MRT buffers.

4. Audio Layers, Procedural SFX & Reactive Signals

  • Soundtrack layers — .audio media nodes with frame-exact lifecycle management.
  • Procedural SFX — deterministic CPU-synthesized whooshes, impacts, risers, downshifters, and glitches positioned on the bar/beat grid (renderSfx, mixSfxInto, softLimit).
  • Multi-track mixing — grasp tracks headlessly with combineAudioTracks and encodeStereoWav.
  • Reactive indicators — drive transforms, scale, borders, or shader uniforms from tempo indicators (Signal.builder) or audio evaluation.

Layer.chart builds line, space, bar (grouped or stacked), scatter, candlestick, pie and donut charts. d3 computes the scales, ticks and geometry; each bar, line, slice and label is an odd field, path or textual content node:

  • Labels use the composition’s registered fonts and the identical GPU textual content renderer as the remainder of the movie.
  • A built-in reveal attracts traces on, grows bars from the baseline and staggers factors and slices (animate: { begin, length, stagger, ease }, or animate: false).
  • The chart is one field, so it positions, animates, grades and tilts into 3D like another layer.
Layer.chart(
  {
    sort: "bar",
    width: 900,
    top: 480,
    classes: ["Q1", "Q2", "Q3", "Q4"],
    collection: [
      { name: "Revenue", data: [12, 19, 24, 31] },
      { title: "Costs", knowledge: [8, 11, 13, 15] },
    ],
    yAxis: { format: "$,.0f" },
    animate: { begin: 10, length: 30 },
  },
  { place: "absolute", x: 120, y: 200 },
);

6. On-Device Vision & Tracking

@gitframes/vision runs ONNX fashions through onnxruntime-node (CPU) or onnxruntime-web (WebGPU) and wires each end result into the identical reactive sign floor the remainder of Gitframes consumes.

Tip

Lazy by development. VisionRunner.create(), comp.withVision(...) and VisionNode.connect(...) carry out zero I/O — no downloads, no periods, no file probes. A mannequin is fetched the primary time a job truly runs. To heat up forward of time, name await runner.preload(["detect", "pose"]) (or await imaginative and prescient.prepared() on an hooked up node).

Every mannequin is Apache-2.0, pinned to an immutable Hugging Face revision, and verified by SHA-256 after obtain.

Task Option Model Output
Detect allowDetection RTMDet-Ins t/s/m (OpenMMLab) COCO-80 bins + scores, tracked over time
Segment allowSegmentation RTMDet-Ins (similar ahead cross as detect) Soft per-instance masks, frame-aligned
Pose allowPose RTMO t/s/m (OpenMMLab) 17 COCO keypoints + visibility per particular person
Matte allowMatte MediaPipe Selfie Segmenter (Google) Fast person-vs-background alpha for portrait / webcam framing

  • Variants — variant: "t" | "s" | "m" (default "s"; ~24 / 43 / 116 MB for RTMDet-Ins). Tune confidence and a COCO courses filter per composition. On CPU, a 2K body takes roughly 200–340 ms to detect + section, ~120 ms for pose and ~20 ms for the matte with "s".
  • One cross, two duties — detection and segmentation share a single RTMDet-Ins inference per body.
  • Picking a matte — the Selfie Segmenter is tuned for an individual filling a lot of the body: it misses distant figures and might report “particular person” on close-ups with no person in them. For the rest, minimize out with occasion masks (matteSource: "occasion", the default).
  • Whole-subject cutouts — masks / matte / crop modes merge each comparably sized occasion that overlaps the primary topic, so a flowing costume or a held instrument stays hooked up to the particular person, whereas a tunnel or window framing them doesn’t.
  • One-frame delay — imaginative and prescient reads every layer’s earlier rendered body, so outcomes path the plate by one body and body 0 has none. Verify imaginative and prescient layers with the exported video or consecutive frames, not body grids.
  • Model cache & mirrors — fashions are cached atomically (temp + rename) in $GITFRAMES_MODELS_DIR (default ~/.cache/gitframes/fashions). Point baseUrl or GITFRAMES_MODELS_BASE_URL at your individual mirror for air-gapped or CI renders.

Temporal monitoring & evaluation

  • Multi-object tracker (TemporalObjectTracker) assigns steady trackIds through IoU affiliation, with configurable minHits, positionSmoothing, and velocity-based coasting for as much as maxMissedFrames (default 15) so a transient miss holds the observe as an alternative of flashing.
  • Pose↔observe matching (pose-track-matcher) binds keypoints to the proper observe by id, then by spatial IoU fallback.
  • One-shot sequence evaluation — comp.analyzeVisionSequence(src, { duties, classes }) decodes frames via the mediabunny pipeline, tracks them, and returns a zod-serializable report (per-track body ranges, imply velocity, sampled heart paths, per-class presence/confidence, imply masks protection, mannequin obtain bytes/timing) (analyzeSequence).

Every tracked entity is uncovered as reactive ProgrammaticSignals that animate layers and shader uniforms:

Group Highlights
objects get(trackId), byCategory(cat, rank), major, depend, hasCategory, detectedCategories
objects.*.bounds x/y/width/top, screenX/screenY/screenWidth/screenHeight, aspectRatio, space
objects.*.anchors 9 anchors (corners, edges, heart) prepared for pinning
objects.*.kinematics vx, vy, velocity, acceleration, headingRad/Deg
objects.*.pose All 17 COCO keypoints, plus hasPose, wristSpeed, handRaised, bodyTiltAngle
masks get(trackId), topic, depend; per-mask space, protection, solidity, bboxFill
segmentation topic, humanSilhouette, occasionMasks, matte.protection, GPU stencilTexture
courses Per-class depend, maxConfidence, current, major, plus a detection histogram
Tensors poseLandmarksTensor [17,3], objectsTensor [16,8], masksTensor [16,2], histogramTensor [80]

Project normalized landmarks to display area with a configurable digicam FOV, then bind any node to a observe or landmark (SpatialLandmarkTransformer, spatial-pin):

  • pinToObject(observe, { anchor, offsetX/Y/Z, matchWidth, matchHeight, cleanFrames, concealWhenMisplaced })
  • pinToLandmark(coord, { offsetX/Y/Z })

High-level composition helpers

  • Subject Sandwich — comp.addSubjectSandwich({ supply, behind, feather, match }) cuts the foreground topic out and locations typography/graphics behind them.
  • Smart Reframing — comp.addSmartFraming({ supply, goal, targetAspect, damping, leadHeadroom }) auto-crops 16:9 → 9:16 whereas monitoring goal.
  • Subject Outline — comp.addSubjectOutline(imaginative and prescient.segmentation.topic, { supply, shade, width, blur }) strokes the segmented boundary as an audio-reactive contour glow.
  • Tracked Region Blur — layer.blurRegion(observe, { power }) blurs faces, plates, or any detected class.
  • Node modes — passthrough, masks, matte, crop, skeleton, bins, monitoring; decide the cutout alpha with matteSource: "occasion" | "selfie", and optionally keyBackground to develop the topic into related foreground.
  • Runtime config is zod-validated and obtainable from a zod-only entry (@gitframes/imaginative and prescient/schemas) so the new path stays zod-free. Unknown or eliminated choices are rejected, not silently ignored.
  • imaginative and prescient.abstract(body) returns a deterministic, serializable snapshot (objects, courses, masks) protected to name inside a body hook.
  • Clear failures — a mannequin that’s the unsuitable dimension, fails its checksum, or lacks an anticipated output raises an error naming the mannequin and its supply.
  • Browser entry — @gitframes/imaginative and prescient/internet re-exports the engine plus createWebGPUProvider() / hasWebGPU(); onnxruntime-web is an elective lazy peer.

7. Headless Conformance & FrameGrid Testing

  • Pixel-sampling invariant assertions — check compositions in Vitest with skia-canvas to confirm shader math, font protection, and Mean Squared Error (MSE) temporal deltas.
  • FrameGrid contact sheets — comp.renderFrameGrid(...) outputs sequential-frame contact sheets for immediate evaluate of easing, kinetic sort, and transitions.

8. Live Preview within the Browser

  • Runs your composition, not a video — beginPreview({ entry, export }) serves a localhost WebGPU participant that hundreds the composition’s personal module and renders each body reside within the browser. Nothing is streamed: the server solely fingers over the bundle, the venture’s property, and the soundtrack blended by the export engine.
  • Timeline, waveform & body stepping — play/pause, scrub, step body by body, and skim decision, FPS, length, and audio standing at a look.
  • One steady URL per venture — the port is derived from the working listing, so re-running the preview replaces the working server and any open tab reloads into the brand new model by itself. Close the tab and the server shuts down about 5 seconds later.
  • Shown the place you’re — beginPreview serves the web page and returns its URL as an alternative of opening a browser, so an agent can present it in its personal pane (Claude Code, Codex); cross open: true to open the system browser.
import { beginPreview } from "gitframes";

const session = await beginPreview(
  { entry: new URL("./movie.ts", import.meta.url), export: "constructFilm" },
  { title: "gitframes movie" },
);
console.log(`Preview at ${session.url}`);
await session.closed; // serves till its tab closes or a more moderen preview takes over

gitframes live preview player: WebGPU rendering, waveform timeline, frame stepping, and audio status at 127.0.0.1:41133


Managed with pnpm workspaces and turbo:

gitframes/
├── packages/
│   ├── gitframes/              # Unified SDK (Composition, Layer, LayerAnimation, Signal, results)
│   ├── core/                   # Core AST, Effect base class, DigitalMediaData, imaginative and prescient sorts
│   ├── compositions/           # Layout engine, Flex/Box AST compiler, timeline evaluator
│   ├── webgpu-renderers/       # WGSL shaders, Slug textual content engine, 3D renderer, digicam, lights, supplies
│   ├── tensor-webgpu/          # WebGPU compute pipelines (Canny, depth-to-normals, stream, deflicker, landmarks)
│   ├── imaginative and prescient/                 # ONNX imaginative and prescient engine: detect, section, pose, matte, monitoring, indicators
│   ├── renderer/               # Headless Node.js WebGPU renderer through Dawn, NetCodecs, skia-canvas
│   ├── renderers/              # Higher-level render orchestration
│   ├── media/                  # Media decoding / encoding adapters
│   ├── node-sdk/               # Node renderer contracts and end result schemas
│   ├── server-utils/           # Server transport systems, storage, asset caches
│   └── client-utils/           # Shared browser utilities
├── nodes/                      # 58+ specialised area nodes (VFX, audio, format, node-vision)
├── apps/
│   └── renderer-service/       # Production HTTP / gRPC rendering microservice container
├── examples/                   # Reference compositions and movies
├── plugins/gitframes/          # Agent plugin: expertise solely (setup, compose, results, render)
└── scripts/                    # Build, launch, and plugin validation tooling

Requirements: Node.js ≥ 22. Gitframes makes use of native GPU acceleration through Dawn / WebGPU or Vulkan.


1. Basic Composition & Kinetic Auto-Layout

import { Composition, Layer, LayerAnimation } from "gitframes";

// 1. Initialize a 1080p60 composition
const comp = new Composition({
  width: 1920,
  top: 1080,
  fps: 60,
  lengthFrames: 180, // 3 seconds
  backgroundColor: "#090a0f",
  fonts: ["assets/fonts/Inter.ttf", "assets/fonts/SpaceGrotesk.ttf"],
});

// 2. Define bodily snap-overshoot animations
const cardEntrance = LayerAnimation.create()
  .fadeIn(0, 20, "power2.out")
  .fromTo("y", 60, 0, { begin: 0, finish: 35, ease: "again.out(1.5)" })
  .fromTo("scale", 0.92, 1.0, { begin: 0, finish: 35, ease: "again.out(1.2)" });

// 3. Assemble a responsive flex-layout card
const heroCard = Layer.field({
  width: 720,
  top: 380,
  background: "#141721",
  borderRadius: 24,
  borderColor: "#262b3d",
  borderWidth: 1.5,
  padding: 32,
  kids: [
    Layer.flex({
      dir: "column",
      gap: 16,
      children: [
        Layer.text("GITFRAMES ENGINE", {
          fontSize: 16,
          fontWeight: 700,
          fill: "#6366f1",
          letterSpacing: 2.0,
        }),
        Layer.text("Next-Gen WebGPU Motion", {
          fontSize: 48,
          fontWeight: 700,
          fill: "#f8fafc",
          fontFamily: "SpaceGrotesk",
        }),
        Layer.text("Direct hardware video composition without headless browser overhead.", {
          fontSize: 20,
          fill: "#94a3b8",
          lineHeight: 28,
        }),
      ],
    }),
  ],
}).animate(cardEntrance);

comp.add(heroCard);

2. Unified 3D Scene with Camera & 3D Model

import { Composition, Layer, Layer3D, CameraAnimation, Light } from "gitframes";

const comp = new Composition({ width: 1920, top: 1080, fps: 60, lengthFrames: 300 });

// 1. LookAt 3D digicam with a steady orbit
const digicamAnim = CameraAnimation.digicam().orbit({
  azimuth: { from: -30, to: 30 },
  elevation: { from: 15, to: 15 },
  radius: { to: 1200 },
  begin: 0,
  finish: 300,
});

comp.add(
  Layer.digicam({ x: 960, y: 540, z: -1000, targetX: 960, targetY: 540, targetZ: 0 }).animate(digicamAnim)
);

// 2. Studio lighting
comp.add(Light.ambient("#ffffff", 0.4));
comp.add(Light.directional({ shade: "#e0e7ff", depth: 1.2, x: 500, y: -800, z: -600 }));

// 3. 3D mannequin with skeletal animation
comp.add(
  Layer.glb("property/fashions/character.glb", {
    x: 960,
    y: 640,
    z: 0,
    scale: 2.5,
    materials: "lit",
    loop: true,
  })
);

// 4. 3D prism format carousel
comp.add(
  Layer3D.carousel({
    radius: 400,
    objects: [
      Layer.box({ width: 280, height: 180, background: "#1e293b", borderRadius: 16 }),
      Layer.box({ width: 280, height: 180, background: "#334155", borderRadius: 16 }),
      Layer.box({ width: 280, height: 180, background: "#0f172a", borderRadius: 16 }),
    ],
  })
);

3. Audio Soundtrack, Procedural SFX & Reactive Signals

import { Composition, Layer, LayerAnimation, Signal, renderSfx, mixSfxInto, softLimit } from "gitframes";

const comp = new Composition({ width: 1920, top: 1080, fps: 60 });
const wholeFrames = 240;

// 1. Soundtrack layer
comp.addAudio(Layer.audio("property/rating.mp3", { quantity: 0.9, lengthFrames: wholeFrames }));

// 2. Frame-accurate procedural SFX on the beat grid
const mattress: [Float32Array, Float32Array] = [
  new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
  new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
];
mixSfxInto(mattress, [
  renderSfx({ type: "whoosh", atBar: 0.79, volume: 0.5 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 1 }),
  renderSfx({ type: "impact", atBar: 1.0, volume: 0.8 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 2 }),
]);
softLimit(mattress);

// 3. Tempo sign (120 BPM = 2 Hz)
const beatPulse = Signal.builder({ sort: "sawtooth", frequency: 2, amplitude: 0.08, offset: 1.0 });

// 4. Bind it to visuals
const reactiveCard = Layer.field({ width: 400, top: 250, background: "#1c202e", borderRadius: 20 })
  .animate(
    LayerAnimation.create()
      .sign("scale", beatPulse, { multiplier: 1.0, offset: 0.0 })
      .fromTo("opacity", 0, 1, { begin: 0, finish: 15, ease: "power2.out" }),
  );

comp.add(reactiveCard);

4. Chained WebGPU Post-Processing VFX

import { Composition, FilmGrain, Vignette, ColorSteadiness } from "gitframes";

const comp = new Composition({ width: 1920, top: 1080, fps: 60 });

// Whole-composition cinematic grade + movie emulsion
comp.apply(new Vignette({ power: 0.28, radius: 0.85 }));
comp.apply(new FilmGrain({ power: 0.06, dimension: 1.5, animated: true }));
comp.apply(
  new ColorSteadiness({
    shadows: { cyanRed: 0, magentaGreen: 2, yellowBlue: 6 },
    highlights: { cyanRed: 4, magentaGreen: 1, yellowBlue: -2 },
  }),
);

5. Vision: Pin, Matte & Reframe

import { Composition, Layer, Vignette } from "gitframes";

const comp = new Composition({ width: 1920, top: 1080, fps: 30 });

// Run imaginative and prescient on the entire composition. Models obtain lazily on first use.
const imaginative and prescient = comp.withVision({
  allowDetection: true,
  allowSegmentation: true,
  allowPose: true,
  variant: "s",
  confidence: 0.35,
});

// Pin a caption to the first tracked topic (smoothing + auto-hide when misplaced)
comp.add(
  Layer.textual content("SUBJECT 01", { fontSize: 40, fill: "#f8fafc" }).pinToObject(
    imaginative and prescient.objects.major,
    { anchor: "topCenter", offsetY: -48, cleanFrames: 5, concealWhenMisplaced: true },
  ),
);

// Drive a shader uniform from a reactive sign — right here, topic masks protection
comp.add(
  Layer.field({ width: 1920, top: 1080, background: "#000000" }).withEffect(
    new Vignette({ power: imaginative and prescient.segmentation.topic.protection, radius: 0.9 }),
  ),
);

// Or use the one-liners for the frequent editorial strikes:
// comp.addSubjectSandwich({ supply: "property/dancer.mp4", behind: [headline], feather: 4 });
// comp.addSmartFraming({ supply: "property/motion.mp4", goal: imaginative and prescient.objects.major, targetAspect: 9 / 16 });
// comp.addSubjectOutline(imaginative and prescient.segmentation.topic, { supply: "property/character.mp4", shade: "#FF5A1F", width: 6 });

// Inspect a supply earlier than authoring: one-shot, ffmpeg-free, zod-serializable report
const report = await comp.analyzeVisionSequence("property/road.mp4", {
  duties: ["detect", "pose"],
  classes: ["person"],
});
console.log(report.tracks.map((t) => `${t.class}#${t.trackId} ${t.frames.be a part of("–")}`));

Standalone runner (no composition):

import { VisionRunner } from "@gitframes/imaginative and prescient";

const runner = VisionRunner.create({ variant: "s", confidence: 0.3 }); // zero I/O
const body = { knowledge: rgba, width: 1920, top: 1080 };
const bins = await runner.detect(body); // downloads RTMDet-Ins on first name
const { masks } = await runner.section(body); // similar ahead cross, no second inference
const { folks } = await runner.pose(body); // RTMO, COCO-17 keypoints
runner.shut();

In the browser (WebGPU EP):

import { VisionRunner, createWebGPUProvider, hasWebGPU } from "@gitframes/imaginative and prescient/internet";

if (hasWebGPU()) {
  const runner = VisionRunner.create({ supplier: createWebGPUProvider() });
}

6. Headless Video & FrameGrid Rendering

import { buildMyComposition } from "./my-composition.js";

const comp = await buildMyComposition();

// 1. Single body to a PNG buffer for visible inspection
const frameBuffer = await comp.renderFrame({ body: 45 });

// 2. Contact-sheet grid of 12 sequential frames
const gridBuffer = await comp.renderFrameGrid({
  beginFrame: 0,
  finishFrame: 120,
  stepFrames: 10,
  cellWidth: 320,
  presentLabels: true,
});

// 3. Final hardware-encoded MP4 with blended audio
const { filePath } = await comp.renderVideo({
  outputPath: "output/final-product-film.mp4",
  high quality: "excessive",
  concurrency: 4,
});

console.log(`Video rendered efficiently to: ${filePath}`);

Engineering Doctrines & Best Practices

  1. Design tokens & theme contracts — outline a centralized THEME for colours, sort, radii, and spacing. Never hardcode magic hex values or ad-hoc margins.
  2. WebGPU premultiplied-alpha invariant — fragment shaders outputting premultiplied alpha (shade * opacity * alpha) should use srcFactor: "one" of their mix state ({ srcFactor: "one", dstFactor: "one-minus-src-alpha", operation: "add" }). Never use srcFactor: "src-alpha" for premultiplied output — squaring alpha darkens fades into murky grey.
  3. Carrier match cuts — carry a visible ingredient (badge, card, cursor, container) throughout scene boundaries with steady velocity and place to keep away from jarring cuts.
  4. Physical easing vocabulary — again.out(1.4–1.7) for snap-overshoot entrances, spring / expo.out for decelerating movement, power2.in for exits. Reserve linear for infinite spinners and time counters.
  5. Headless invariant verification — confirm shader transforms, glyph protection, and temporal MSE deltas with skia-canvas pixel sampling in Vitest earlier than delivery.

Gitframes ships agent expertise that educate Claude, Codex, and different coding brokers learn how to write, render, and examine compositions. The plugin (gitframes) is listed in Anthropic’s official plugin listing and comprises solely expertise — no MCP servers, hooks, or instructions. Every different agent will get the identical expertise via the skills CLI.

Skill Use it for
gitframes Starting a venture: set up from npm, scaffold a composition and render script, first verified render
gitframes-compose Compositions, layer bushes, format, animation and easing, beat grids, movie construction
gitframes-effects Effect courses, the unified part structure, premultiplied-alpha invariants, imaginative and prescient conditioning
gitframes-render Headless rendering, FrameGrid inspection, pixel probes, MP4 supply checks

Once put in, expertise load routinely when a job matches (e.g. “add a film-grain cross to this scene” or “render a body grid of intro.ts”).

What the plugin runs and sends

The plugin is directions solely. It bundles no executables, MCP servers, hooks, or bundle launchers, and it sends no knowledge wherever. The expertise inform your agent so as to add the gitframes npm bundle to your venture and learn how to use it. When that code makes use of on-device imaginative and prescient, the SDK downloads the pinned mannequin weights from Hugging Face on first use (see On-Device Vision). Nothing else leaves your machine.

/plugin set up gitframes

Or out of your shell:

claude plugin set up gitframes@claude-plugins-official

It installs from Anthropic’s official market, which Claude Code provides for you, so there isn’t a market step, and plugins from it replace routinely. Afterwards, restart Claude Code or run /reload-plugins. /plugin instructions want an interactive claude terminal; within the desktop app’s Code tab, use the shell type or + > Plugins > Add plugin and decide Gitframes.

Add --scope venture to the shell type to file the plugin in .claude/settings.json for the entire staff.

Enable it for everybody in your repo. Commit this to .claude/settings.json; Claude Code prompts teammates to put in it after they belief the folder:

{
  "enabledPlugins": {
    "gitframes@claude-plugins-official": true
  }
}

Straight from this repository (tracks foremost as an alternative of the listing launch):

/plugin market add gatewai-dev/gitframes
/plugin set up gitframes@gitframes-plugins

Codex, Cursor, Hermes, and different brokers

The skills CLI installs the talents into any of 70+ brokers, together with Codex, Cursor, Hermes, Gemini CLI, GitHub Copilot, Windsurf, OpenCode, and Goose:

npx expertise add gatewai-dev/gitframes

It detects the brokers in your machine and asks the place to put in. To select them your self, cross -a as soon as per agent, add -g to put in in your consumer as an alternative of this venture, and -y to skip the prompts:

npx expertise add gatewai-dev/gitframes -a codex -a cursor -a hermes-agent -g -y

Keep them present with npx expertise replace, and take away them with npx expertise take away.

Or copy the folders by hand: put plugins/gitframes/expertise// into .claude/expertise/, .brokers/expertise/, or ~/.brokers/expertise/. VS Code / Copilot / Cursor / Kiro can load the transportable plugin.json via their plugin UI.

The plugin lives in plugins/gitframes/ so installs carry solely the talents; customers get the engine from npm. Two manifests there describe it: plugin.json (transportable Agent Plugins 1.0, which additionally carries the OpenAI itemizing metadata) and .claude-plugin/plugin.json. The market catalog is .claude-plugin/marketplace.json. The transportable discipline set is closed — client-specific fields go in that consumer’s manifest, not in plugin.json. The model in each follows the gitframes bundle: pnpm run model:packages syncs it after changeset model (or run pnpm run sync:plugin-version by itself), since shoppers use it to resolve when to replace.

Inside this repository, Claude Code and different brokers decide up expertise via the symlinks in .brokers/expertise/ and .claude/expertise/. Skills reside solely below plugins/gitframes/expertise/; by no means copy them elsewhere. pnpm run examine:plugins validates manifests, ability frontmatter, market catalogs, symlinks, and the generated results catalog. pnpm run sync:effects-catalog regenerates the gitframes-effects catalog after any Effect class change.


Reference Showcase Examples

The examples/ listing holds production-grade reference compositions:


Gitframes makes use of pnpm (10+) and turbo for orchestration.

# Install
pnpm set up

# Build all packages
pnpm construct

# Run conformance checks
pnpm check

# Check the imaginative and prescient fashions finish to finish (downloads ~380 MB of weights as soon as)
pnpm --filter @gitframes/imaginative and prescient check:fashions

# Render a selected showcase instance
pnpm --filter @gitframes/example-21-full-circle render

# Render the grasp model movie
cd examples/19_gitframes_film && pnpm render

Docker Container for Production Rendering

An optimized Dockerfile.renderer deploys the renderer service into cloud GPU clusters:

docker construct -t gitframes-renderer -f Dockerfile.renderer .


Gitframes is open-source software program licensed below Apache-2.0. The imaginative and prescient fashions it downloads on demand — RTMDet-Ins and RTMO (OpenMMLab) and the Selfie Segmenter (Google) — are additionally Apache-2.0; see registry.ts for precise sources and checksums.



Source link