- ResonateX T1
- ResonateX T1
ResonateX T1
A compact conversational language model trained from scratch
125.86M Parameters Β· 1.000B Training Tokens Β· 18 Transformer Blocks Β· 2048 Context
Trained from random initialization on an NVIDIA A100.
Overview
ResonateX T1 is a 125.86M-parameter decoder-only Transformer language model developed by ResonateX Integrated Technologies.
T1 was trained from random initialization. It is not a fine-tune of an existing pretrained language model, and no pretrained model weights were used to initialize it.
The primary objective of T1 is deliberately simple:
Learn to speak before trying to know everything.
The model focuses on English language generation and conversational behavior: constructing coherent sentences, maintaining short conversational context, producing assistant-style responses, and serving as a compact foundation for further experimentation.
At approximately 126M parameters, T1 is small enough to train, inspect, modify, continue pretraining, quantize, and experiment with without billion-parameter-scale hardware.
ResonateX T1 is an experimental research language model. It has not yet undergone comprehensive standardized benchmark evaluation or production safety validation.
Model at a Glance
| Property | ResonateX T1 |
|---|---|
| Parameters | 125,857,536 |
| Architecture | Decoder-only Transformer |
| Transformer Blocks | 18 |
| Hidden Size | 768 |
| FFN Size | 2,048 |
| Query Heads | 12 |
| KV Heads | 4 |
| Head Dimension | 64 |
| Attention | Grouped-Query Attention |
| Position Encoding | RoPE |
| Normalization | RMSNorm |
| FFN Activation | SwiGLU |
| Vocabulary | 16,384 |
| Context Length | 2,048 tokens |
| Training Tokens | 1,000,046,592 |
| Optimizer Steps | 10,173 |
| Precision | BF16 |
| Primary Language | English |
| Objective | Causal Language Modeling |
| Weights | SafeTensors |
Architecture
ResonateX T1 uses a compact decoder-only Transformer architecture built around RMSNorm, RoPE, Grouped-Query Attention, SwiGLU, and tied input/output embeddings.
ResonateX T1
β
βββ Parameters .............. 125,857,536
βββ Transformer Blocks ...... 18
βββ Hidden Size ............. 768
βββ FFN Size ................ 2,048
β
βββ Attention
β βββ Query Heads ......... 12
β βββ KV Heads ............ 4
β βββ Head Dimension ...... 64
β βββ Type ................ Grouped-Query Attention
β
βββ Position Encoding ....... RoPE
βββ Normalization ........... RMSNorm
βββ Feed Forward ............ SwiGLU
βββ Vocabulary .............. 16,384
βββ Context ................. 2,048
βββ Embedding / LM Head ..... Tied
Grouped-Query Attention
T1 uses 12 query heads and 4 key/value heads.
Grouped-Query Attention reduces the number of independent key/value representations while retaining multiple query heads.
Rotary Position Embeddings
RoPE is used for positional information inside the attention mechanism.
SwiGLU
Transformer feed-forward blocks use SwiGLU with an intermediate dimension of 2,048.
Tied Embeddings
The token embedding matrix and language-model output projection share weights.
Trained From Scratch
ResonateX T1 began with randomly initialized weights.
Pretrained model weights NO
Fine-tuned base model NO
Distillation NO
Random initialization YES
Original training run YES
No existing pretrained language model was used as the starting checkpoint.
Tokenizer
ResonateX T1 uses its own tokenizer prepared specifically for the model.
Vocabulary Size 16,384
Context Length 2,048
The conversational format uses dedicated special tokens:
<|system|>
<|user|>
<|assistant|>
<|end|>
The inference prompt also uses <|bos|> to begin the sequence.
Training
The completed training run processed slightly more than one billion tokens.
ββββββββββββββββββββββββββββββββββββββββββββ
β RESONATEX T1 β TRAINING RUN β
β βββββββββββββββββββββββββββββββββββββββββββ£
β Parameters 125,857,536 β
β Tokens Seen 1,000,046,592 β
β Optimizer Steps 10,173 β
β Tokens / Update 98,304 β
β Precision BF16 β
β GPU A100-SXM4-40GB β
β Training Time 102.75 min β
β Throughput 162,214 tok/s β
ββββββββββββββββββββββββββββββββββββββββββββ
Training was performed on an NVIDIA A100-SXM4-40GB using Google Colab.
Training Configuration
| Setting | Value |
|---|---|
| Sequence Length | 2,048 |
| Physical Batch Size | 24 |
| Gradient Accumulation | 2 |
| Effective Tokens / Update | 98,304 |
| Optimizer | AdamW |
| Adam Betas | (0.9, 0.95) |
| Adam Epsilon | 1e-8 |
| Weight Decay | 0.1 |
| Peak Learning Rate | 6e-4 |
| Minimum Learning Rate | 6e-5 |
| LR Schedule | Warmup + Cosine |
| Warmup | 20M tokens |
| Gradient Clipping | 1.0 |
| Precision | BF16 |
Training Data
T1 was trained using two primary data streams: general English text and conversational data.
| Source | Tokens | Share |
|---|---|---|
| General | 654,262,272 | 65.4% |
| Dialogue | 345,784,320 | 34.6% |
| Total | 1,000,046,592 | 100% |
Progressive Dialogue Curriculum
Dialogue exposure increased during training:
TRAINING START TRAINING END
β β
βΌ βΌ
~25% dialogue ββββββββββββββββββββββββββββΆ ~45% dialogue
Training Result
TRAINING FINISHED
Model ResonateX T1
Parameters 125,857,536
Steps 10,173
Tokens Seen 1,000,046,592
General Tokens 654,262,272
Dialogue Tokens 345,784,320
Elapsed 102.75 minutes
Average Throughput 162,214 tok/s
Target Reached TRUE
Final Training Loss 2.472
Validation
| Split | Loss | Perplexity |
|---|---|---|
| General | 3.08996 |
21.976 |
| Dialogue | 2.57932 |
13.188 |
| Combined | 2.91123 |
β |
Best combined validation loss:
2.9112325310707092
These values describe the held-out validation data used by the training pipeline and are not standardized external benchmark scores.
Design Goal
ResonateX T1 is not designed around the assumption that a 125M model should memorize enormous amounts of factual knowledge.
Its central experiment is:
Can a small language model learn to speak coherently from scratch?
The training emphasizes language structure, conversation, short-context coherence, and generation.
Knowledge retrieval, RAG, tools, domain specialization, and additional training can be layered on top of this foundation.
Quick Start
Install the runtime dependencies:
pip install -U torch transformers safetensors tokenizers
Because ResonateX T1 uses a custom Transformers architecture, load it with trust_remote_code=True.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "ResonateXIntegratedTechnologies/ResonateX-T1"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
torch.bfloat16
if device == "cuda" and torch.cuda.is_bf16_supported()
else torch.float32
)
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=dtype,
).to(device).eval()
The trained parameters are stored in model.safetensors.
The repository-level Python files do not contain the trained model weights:
configuration_resonatex.pydefines the custom configuration.modeling_resonatex.pydefines the ResonateX T1 architecture.
They are loaded automatically by Hugging Face Transformers and do not need to be executed manually.
Prompt Format
User / Assistant
<|bos|><|user|>
Hello!<|end|>
<|assistant|>
System / User / Assistant
<|bos|><|system|>
You are a helpful assistant.<|end|>
<|user|>
Hello!<|end|>
<|assistant|>
The system section is optional.
Generation
Suggested starting parameters:
| Parameter | Value |
|---|---|
| Temperature | 0.8 |
| Top P | 0.92 |
| Top K | 50 |
| Repetition Penalty | 1.08 |
| Max New Tokens | 180 |
Example:
prompt = """<|bos|><|user|>
Hello! Who are you?<|end|>
<|assistant|>
"""
inputs = tokenizer(
prompt,
return_tensors="pt",
add_special_tokens=False,
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=180,
do_sample=True,
temperature=0.8,
top_p=0.92,
top_k=50,
repetition_penalty=1.08,
)
generated = outputs[0, inputs["input_ids"].shape[1]:]
print(
tokenizer.decode(
generated,
skip_special_tokens=True,
)
)
These settings are starting points rather than benchmark-optimized generation parameters.
Example Generations
The following examples are from the model after the 1B-token training run.
Conversation
Prompt
Hello! Who are you?
ResonateX T1
Hello! Iβm an AI language model created by OpenAI, designed to assist with information, answer questions, and engage in conversation. How can I help you today?
The reference to OpenAI is incorrect. It is a learned textual pattern rather than factual self-knowledge.
Simple Instruction
Prompt
Explain why the sky looks blue in simple English.
ResonateX T1
The sky looks blue in simple English.
This demonstrates the distinction between producing language and possessing sufficient factual or instruction-following capability.
Free Generation
The term "space" is a term that means something beyond our current understanding, even if it's not immediately obvious. It suggests that there might be some kind of strange or unusual or unique way we experience space. The term itself is quite intriguing and could be used to describe a variety of things, depending on the context.
The passage maintains grammatical structure and a consistent topic despite limited factual grounding.
Evaluation
ResonateX T1 has not yet undergone a comprehensive standardized benchmark suite.
No official scores are currently claimed for:
MMLU
HellaSwag
ARC
TruthfulQA
GSM8K
HumanEval
MT-Bench
Future releases may include reproducible standardized evaluations.
Until then, ResonateX T1 should be treated as an experimental compact language model rather than a benchmark-validated general-purpose assistant.
Limitations
ResonateX T1 contains 125.86M parameters and was trained on approximately 1B tokens.
Expected limitations include:
- limited factual knowledge
- factual hallucinations
- weak complex reasoning
- weak mathematical reasoning
- limited programming capability
- inconsistent instruction following
- possible repetition
- generic responses
- sensitivity to prompt formatting
- sensitivity to sampling parameters
- degradation during long generations
- limited multilingual capability
The model was trained primarily for English.
The trained context length is 2,048 tokens.
Safety
ResonateX T1 has not undergone comprehensive safety alignment, red-team evaluation, or production safety validation.
Generated text may contain inaccurate, biased, inappropriate, unexpected, or unsafe content.
The model should not be treated as an authoritative source for medical, legal, financial, safety-critical, or similarly high-stakes decisions.
Applications built on the model should implement safeguards appropriate to their use case.
Model Files
The public Hugging Face repository is intentionally compact:
ResonateX-T1/
β
βββ README.md
βββ image.png
β
βββ model.safetensors
βββ config.json
βββ generation_config.json
β
βββ tokenizer.json
βββ tokenizer_config.json
βββ special_tokens_map.json
β
βββ configuration_resonatex.py
βββ modeling_resonatex.py
What each file does
| File | Purpose |
|---|---|
README.md |
Model card and usage documentation |
image.png |
Repository header image |
model.safetensors |
Trained BF16 model parameters |
config.json |
Hugging Face model configuration |
generation_config.json |
Default text-generation settings |
tokenizer.json |
Tokenizer model |
tokenizer_config.json |
Tokenizer configuration |
special_tokens_map.json |
Special-token mapping |
configuration_resonatex.py |
Custom PretrainedConfig implementation |
modeling_resonatex.py |
Custom PreTrainedModel implementation |
model.safetensors contains the trained 125,857,536 parameters in BF16 format.
Training scripts, optimizer state, training datasets, metrics logs, and the full training checkpoint are not required for ordinary inference and are intentionally excluded from the public inference repository.
chat.py, when used locally for testing, is also not required in the public Hugging Face repository.
Research Directions
ResonateX T1
β
ββββββββββββββββΌβββββββββββββββ
β β β
βΌ βΌ βΌ
More Pretraining Dialogue Knowledge
β Training Specialization
β β β
ββββββββββββββββΌβββββββββββββββ
β
βΌ
Future T1 variants
Potential future experiments include continued pretraining, conversational specialization, instruction training, RAG, domain specialization, preference optimization, longer-context training, quantization, GGUF conversion, inference optimization, and standardized evaluation.
Reproducibility
Training run signature:
6af088063a156831a8e82178a05b6813cd4780210268a7a3c263ca46eb654975
Release files can additionally be verified using published SHA-256 hashes when provided with a release.
License
ResonateX T1 is released under the Apache License 2.0.
Users should also review applicable licenses and terms of the training datasets when redistributing or commercially deploying derivative models.
Citation
@software{resonatex_t1_2026,
title = {ResonateX T1},
author = {ResonateX Integrated Technologies},
year = {2026},
description = {A 125.86M parameter conversational language model trained from scratch on 1B tokens}
}
- Downloads last month
- 548