GIN AI Research (gin.info.vn) is an applied research laboratory engineering from-scratch 1.58-bit ternary neural models, real-time autonomous multimodal tutoring classrooms, privacy-first healthcare companions, and open-source developer FinOps engines for autonomous AI coding agents.
We solve fundamental engineering challenges: drastic compute footprint reduction, absolute data sovereignty, and aligning autonomous agents with practical human workflows.
Training from-scratch 1.58-bit ternary architectures with Quantization-Aware Training (QAT) and Straight-Through Estimators (STE). Achieving near-lossless machine translation at ~350 tok/s on standard laptop CPUs with a 77.56 MB disk footprint.
Engineering synchronous browser-based classrooms with sub-second voice activity interruption (Barge-in VAD), interleaved pedagogical reasoning, and autonomous DOM tool calling that dynamically drives visual slides, lip syncing, and stroke animations.
Deep lifecycle hooks telemetry and cost analytics built for autonomous coding agent fleets. Quantifying Prompt Caching savings in USD, visualizing tool execution flows via Sankey graphs, and enforcing zero-prompt privacy.
Evidence-informed maternal and child health research. 100% offline-first PWA, clinical growth percentile calculators, CDC milk-stash management, Edinburgh Postnatal Depression Scale (EPDS) screening, and zero third-party tracking.
Every system is an open-source, verifiable, production-grade working implementation.
An autonomous browser-based classroom delivering synchronous 1-on-1 human-grade language tutoring. Features bidirectional streaming audio (16kHz in, 24kHz out), sub-second natural interruption (Barge-in VAD), interleaved pedagogical reasoning stream, and autonomous DOM tool-calling agent (change_slide, highlight_element, mark_error). Over 1,500+ generated lesson illustrations and 28 lip-synced characters.
From-scratch 1.58-bit ternary weight ({-1, 0, +1}) Neural Machine Translation (Japanese ↔ Vietnamese) based on the BitNet b1.58 architecture. 152.1M parameters trained on 15.98M parallel pairs and compressed down to a 77.56 MB GGUF binary. Runs at ~350 tokens/sec on standard laptop CPUs with only ~125 MB working RAM.
| Model / System | Accuracy (Score 2) | Speed (CPU) | Size / RAM |
|---|---|---|---|
| Bit-Translate v7a i2s | 78.5% | ~350 tok/s | 77.56 MB / ~125 MB |
| Google Translate Cloud | 69.0% | Cloud API Latency | Cloud Bound |
A self-hosted enterprise control center built for autonomous coding agent fleets (such as Claude Code and custom agentic CLI harnesses). Hooks directly into lifecycle events (SessionStart, PreToolUse, PostToolUse, Stop). Tracks token burn, computes real USD savings from Prompt Caching, visualizes tool transition Sankeys, and tracks team adoption—all without ever storing prompt text or corporate source code.
SessionStart ➔ ToolExecution ➔ CostEngine
• Token Input / Output Breakdown
• Cache Creation vs. Cache Read Hits
• Tool Latency Percentiles (P50/P90/P99)
• Task Classification (Debug/Code/Plan)
A desktop translation overlay engineered for high-security, defense, and confidential corporate boardrooms. Completely air-gapped: zero external HTTP calls during inference. Combines local Faster-Whisper (CTranslate2 INT8, 2–4x faster than whisper.cpp) with NLLB-200 INT8 (~594 MB) via an English pivot and contextual sliding-window accumulation.
A calm, private, evidence-informed parenting and digital health research companion. 15.6k lines of Dart with 97 passing automated tests. Built specifically for pumping mothers (dual-side timer, CDC milk stash, estimated direct latch). Features drag-to-fill bottle with AAP citations, EASY routine nap engine, WHO 2006 growth percentiles, and Edinburgh Postnatal Depression Scale (EPDS) self-screening. Zero ads, zero tracking, all data stays on device.
We hold our work to strict empirical standards: zero metric extrapolation, reproducible test benches, and physical hardware verification.
Every figure in our papers and repos originates from executed tool runs, physical OS counters, or blinded multi-model audits. Unmeasured benchmarks are strictly labeled as untested.
We design systems where user data never leaves the edge device by default. Our enterprise FinOps and digital health suites enforce strict client-side indexing and zero-telemetry architectures.
All codebases, conversion scripts, synthetic teacher prompts, and testing suites are published publicly on GitHub with complete instructions to reproduce every finding locally.
Target milestones across model scaling, clinical trials, and multi-agent developer infrastructure.
Scale Bit-Translate architecture to a 300M-parameter tier under a 150MB binary budget. Deploy multi-tenant hosted Sensei Classroom for pilot cohorts preparing for JLPT examinations.
Expand KS-Dashboard adapters beyond CLI to multi-agent harnesses (Cursor, Aider, OpenHands). Release automated budget threshold webhooks and circuit breakers.
Conduct an ethical, consented pilot study with new mothers in Vietnam using GinBaby to evaluate the impact of private postpartum depression tracking and feeding routines.
Whether you are a researcher, open-source contributor, academic institution, or technology partner, we welcome inquiries, peer discussions, and collaborative engineering.