ResonateX T1

ResonateX T1

A compact conversational language model trained from scratch

125.86M Parameters Β· 1.000B Training Tokens Β· 18 Transformer Blocks Β· 2048 Context

Parameters Training Tokens Context Precision License

Trained from random initialization on an NVIDIA A100.


Overview

ResonateX T1 is a 125.86M-parameter decoder-only Transformer language model developed by ResonateX Integrated Technologies.

T1 was trained from random initialization. It is not a fine-tune of an existing pretrained language model, and no pretrained model weights were used to initialize it.

The primary objective of T1 is deliberately simple:

Learn to speak before trying to know everything.

The model focuses on English language generation and conversational behavior: constructing coherent sentences, maintaining short conversational context, producing assistant-style responses, and serving as a compact foundation for further experimentation.

At approximately 126M parameters, T1 is small enough to train, inspect, modify, continue pretraining, quantize, and experiment with without billion-parameter-scale hardware.

ResonateX T1 is an experimental research language model. It has not yet undergone comprehensive standardized benchmark evaluation or production safety validation.


Model at a Glance

Property ResonateX T1
Parameters 125,857,536
Architecture Decoder-only Transformer
Transformer Blocks 18
Hidden Size 768
FFN Size 2,048
Query Heads 12
KV Heads 4
Head Dimension 64
Attention Grouped-Query Attention
Position Encoding RoPE
Normalization RMSNorm
FFN Activation SwiGLU
Vocabulary 16,384
Context Length 2,048 tokens
Training Tokens 1,000,046,592
Optimizer Steps 10,173
Precision BF16
Primary Language English
Objective Causal Language Modeling
Weights SafeTensors

Architecture

ResonateX T1 uses a compact decoder-only Transformer architecture built around RMSNorm, RoPE, Grouped-Query Attention, SwiGLU, and tied input/output embeddings.

ResonateX T1
β”‚
β”œβ”€β”€ Parameters .............. 125,857,536
β”œβ”€β”€ Transformer Blocks ...... 18
β”œβ”€β”€ Hidden Size ............. 768
β”œβ”€β”€ FFN Size ................ 2,048
β”‚
β”œβ”€β”€ Attention
β”‚   β”œβ”€β”€ Query Heads ......... 12
β”‚   β”œβ”€β”€ KV Heads ............ 4
β”‚   β”œβ”€β”€ Head Dimension ...... 64
β”‚   └── Type ................ Grouped-Query Attention
β”‚
β”œβ”€β”€ Position Encoding ....... RoPE
β”œβ”€β”€ Normalization ........... RMSNorm
β”œβ”€β”€ Feed Forward ............ SwiGLU
β”œβ”€β”€ Vocabulary .............. 16,384
β”œβ”€β”€ Context ................. 2,048
└── Embedding / LM Head ..... Tied

Grouped-Query Attention

T1 uses 12 query heads and 4 key/value heads.

Grouped-Query Attention reduces the number of independent key/value representations while retaining multiple query heads.

Rotary Position Embeddings

RoPE is used for positional information inside the attention mechanism.

SwiGLU

Transformer feed-forward blocks use SwiGLU with an intermediate dimension of 2,048.

Tied Embeddings

The token embedding matrix and language-model output projection share weights.


Trained From Scratch

ResonateX T1 began with randomly initialized weights.

Pretrained model weights     NO
Fine-tuned base model        NO
Distillation                 NO

Random initialization        YES
Original training run        YES

No existing pretrained language model was used as the starting checkpoint.


Tokenizer

ResonateX T1 uses its own tokenizer prepared specifically for the model.

Vocabulary Size     16,384
Context Length      2,048

The conversational format uses dedicated special tokens:

<|system|>
<|user|>
<|assistant|>
<|end|>

The inference prompt also uses <|bos|> to begin the sequence.


Training

The completed training run processed slightly more than one billion tokens.

╔══════════════════════════════════════════╗
β•‘        RESONATEX T1 β€” TRAINING RUN      β•‘
╠══════════════════════════════════════════╣
β•‘ Parameters          125,857,536          β•‘
β•‘ Tokens Seen       1,000,046,592          β•‘
β•‘ Optimizer Steps          10,173          β•‘
β•‘ Tokens / Update          98,304          β•‘
β•‘ Precision                 BF16           β•‘
β•‘ GPU          A100-SXM4-40GB              β•‘
β•‘ Training Time       102.75 min           β•‘
β•‘ Throughput      162,214 tok/s            β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•

Training was performed on an NVIDIA A100-SXM4-40GB using Google Colab.

Training Configuration

Setting Value
Sequence Length 2,048
Physical Batch Size 24
Gradient Accumulation 2
Effective Tokens / Update 98,304
Optimizer AdamW
Adam Betas (0.9, 0.95)
Adam Epsilon 1e-8
Weight Decay 0.1
Peak Learning Rate 6e-4
Minimum Learning Rate 6e-5
LR Schedule Warmup + Cosine
Warmup 20M tokens
Gradient Clipping 1.0
Precision BF16

Training Data

T1 was trained using two primary data streams: general English text and conversational data.

Source Tokens Share
General 654,262,272 65.4%
Dialogue 345,784,320 34.6%
Total 1,000,046,592 100%

Progressive Dialogue Curriculum

Dialogue exposure increased during training:

TRAINING START                              TRAINING END
     β”‚                                           β”‚
     β–Ό                                           β–Ό
  ~25% dialogue  ━━━━━━━━━━━━━━━━━━━━━━━━━━━▢  ~45% dialogue

Training Result

TRAINING FINISHED

Model                 ResonateX T1
Parameters            125,857,536
Steps                 10,173
Tokens Seen           1,000,046,592

General Tokens        654,262,272
Dialogue Tokens       345,784,320

Elapsed               102.75 minutes
Average Throughput    162,214 tok/s

Target Reached        TRUE
Final Training Loss   2.472

Validation

Split Loss Perplexity
General 3.08996 21.976
Dialogue 2.57932 13.188
Combined 2.91123 β€”

Best combined validation loss:

2.9112325310707092

These values describe the held-out validation data used by the training pipeline and are not standardized external benchmark scores.


Design Goal

ResonateX T1 is not designed around the assumption that a 125M model should memorize enormous amounts of factual knowledge.

Its central experiment is:

Can a small language model learn to speak coherently from scratch?

The training emphasizes language structure, conversation, short-context coherence, and generation.

Knowledge retrieval, RAG, tools, domain specialization, and additional training can be layered on top of this foundation.


Quick Start

Install the runtime dependencies:

pip install -U torch transformers safetensors tokenizers

Because ResonateX T1 uses a custom Transformers architecture, load it with trust_remote_code=True.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "ResonateXIntegratedTechnologies/ResonateX-T1"

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = (
    torch.bfloat16
    if device == "cuda" and torch.cuda.is_bf16_supported()
    else torch.float32
)

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    trust_remote_code=True,
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=dtype,
).to(device).eval()

The trained parameters are stored in model.safetensors.

The repository-level Python files do not contain the trained model weights:

  • configuration_resonatex.py defines the custom configuration.
  • modeling_resonatex.py defines the ResonateX T1 architecture.

They are loaded automatically by Hugging Face Transformers and do not need to be executed manually.


Prompt Format

User / Assistant

<|bos|><|user|>
Hello!<|end|>
<|assistant|>

System / User / Assistant

<|bos|><|system|>
You are a helpful assistant.<|end|>
<|user|>
Hello!<|end|>
<|assistant|>

The system section is optional.


Generation

Suggested starting parameters:

Parameter Value
Temperature 0.8
Top P 0.92
Top K 50
Repetition Penalty 1.08
Max New Tokens 180

Example:

prompt = """<|bos|><|user|>
Hello! Who are you?<|end|>
<|assistant|>
"""

inputs = tokenizer(
    prompt,
    return_tensors="pt",
    add_special_tokens=False,
).to(model.device)

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=180,
        do_sample=True,
        temperature=0.8,
        top_p=0.92,
        top_k=50,
        repetition_penalty=1.08,
    )

generated = outputs[0, inputs["input_ids"].shape[1]:]

print(
    tokenizer.decode(
        generated,
        skip_special_tokens=True,
    )
)

These settings are starting points rather than benchmark-optimized generation parameters.


Example Generations

The following examples are from the model after the 1B-token training run.

Conversation

Prompt

Hello! Who are you?

ResonateX T1

Hello! I’m an AI language model created by OpenAI, designed to assist with information, answer questions, and engage in conversation. How can I help you today?

The reference to OpenAI is incorrect. It is a learned textual pattern rather than factual self-knowledge.

Simple Instruction

Prompt

Explain why the sky looks blue in simple English.

ResonateX T1

The sky looks blue in simple English.

This demonstrates the distinction between producing language and possessing sufficient factual or instruction-following capability.

Free Generation

The term "space" is a term that means something beyond our current understanding, even if it's not immediately obvious. It suggests that there might be some kind of strange or unusual or unique way we experience space. The term itself is quite intriguing and could be used to describe a variety of things, depending on the context.

The passage maintains grammatical structure and a consistent topic despite limited factual grounding.


Evaluation

ResonateX T1 has not yet undergone a comprehensive standardized benchmark suite.

No official scores are currently claimed for:

MMLU
HellaSwag
ARC
TruthfulQA
GSM8K
HumanEval
MT-Bench

Future releases may include reproducible standardized evaluations.

Until then, ResonateX T1 should be treated as an experimental compact language model rather than a benchmark-validated general-purpose assistant.


Limitations

ResonateX T1 contains 125.86M parameters and was trained on approximately 1B tokens.

Expected limitations include:

  • limited factual knowledge
  • factual hallucinations
  • weak complex reasoning
  • weak mathematical reasoning
  • limited programming capability
  • inconsistent instruction following
  • possible repetition
  • generic responses
  • sensitivity to prompt formatting
  • sensitivity to sampling parameters
  • degradation during long generations
  • limited multilingual capability

The model was trained primarily for English.

The trained context length is 2,048 tokens.


Safety

ResonateX T1 has not undergone comprehensive safety alignment, red-team evaluation, or production safety validation.

Generated text may contain inaccurate, biased, inappropriate, unexpected, or unsafe content.

The model should not be treated as an authoritative source for medical, legal, financial, safety-critical, or similarly high-stakes decisions.

Applications built on the model should implement safeguards appropriate to their use case.


Model Files

The public Hugging Face repository is intentionally compact:

ResonateX-T1/
β”‚
β”œβ”€β”€ README.md
β”œβ”€β”€ image.png
β”‚
β”œβ”€β”€ model.safetensors
β”œβ”€β”€ config.json
β”œβ”€β”€ generation_config.json
β”‚
β”œβ”€β”€ tokenizer.json
β”œβ”€β”€ tokenizer_config.json
β”œβ”€β”€ special_tokens_map.json
β”‚
β”œβ”€β”€ configuration_resonatex.py
└── modeling_resonatex.py

What each file does

File Purpose
README.md Model card and usage documentation
image.png Repository header image
model.safetensors Trained BF16 model parameters
config.json Hugging Face model configuration
generation_config.json Default text-generation settings
tokenizer.json Tokenizer model
tokenizer_config.json Tokenizer configuration
special_tokens_map.json Special-token mapping
configuration_resonatex.py Custom PretrainedConfig implementation
modeling_resonatex.py Custom PreTrainedModel implementation

model.safetensors contains the trained 125,857,536 parameters in BF16 format.

Training scripts, optimizer state, training datasets, metrics logs, and the full training checkpoint are not required for ordinary inference and are intentionally excluded from the public inference repository.

chat.py, when used locally for testing, is also not required in the public Hugging Face repository.


Research Directions

                       ResonateX T1
                            β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚              β”‚              β”‚
             β–Ό              β–Ό              β–Ό
       More Pretraining   Dialogue      Knowledge
             β”‚            Training     Specialization
             β”‚              β”‚              β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β”‚
                            β–Ό
                     Future T1 variants

Potential future experiments include continued pretraining, conversational specialization, instruction training, RAG, domain specialization, preference optimization, longer-context training, quantization, GGUF conversion, inference optimization, and standardized evaluation.


Reproducibility

Training run signature:

6af088063a156831a8e82178a05b6813cd4780210268a7a3c263ca46eb654975

Release files can additionally be verified using published SHA-256 hashes when provided with a release.


License

ResonateX T1 is released under the Apache License 2.0.

Users should also review applicable licenses and terms of the training datasets when redistributing or commercially deploying derivative models.


Citation

@software{resonatex_t1_2026,
  title        = {ResonateX T1},
  author       = {ResonateX Integrated Technologies},
  year         = {2026},
  description  = {A 125.86M parameter conversational language model trained from scratch on 1B tokens}
}

ResonateX T1

Trained from scratch. Built to talk.

125.86M parameters Β· 1.000B tokens Β· 10,173 steps Β· 2048 context

ResonateX Integrated Technologies Β· 2026

Downloads last month
548
Safetensors
Model size
0.1B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ResonatexIntegratedTechnologies/ResonateX-T1-125M-Talking

Quantizations
1 model