Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Thilak Ruwan Chamara Denipitiya
ruwan
2
2
105
Follow
21world's profile picture
Quazim0t0's profile picture
webxos's profile picture
3 followers
·
51 following
ThilakDen
AI & ML interests
AI
Recent Activity
reacted
to
FredyRivera-dev
's
post
with 🚀
about 3 hours ago
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen. Covers: - Hybrid architecture based on Qwen3.5 - Pre-training with 15B tokens - Cost benchmark between H200 and B200 - Post-training with SFT + LoRA - Full code and data, open source With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline. Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
liked
a model
about 16 hours ago
nightmedia/Qwen3.5-9B-Holodeck-Fara
liked
a model
about 17 hours ago
microsoft/Fara1.5-9B
View all activity
Organizations
None yet
models
5
Sort: Recently updated
ruwan/MyGemmaNPC
Text Generation
•
0.3B
•
Updated
Aug 24, 2025
•
5
ruwan/story_model_final
16.2M
•
Updated
Mar 25, 2025
•
1
ruwan/llama-tiny
Text Generation
•
Updated
Dec 11, 2023
•
6
ruwan/open-llama-sharded-1GB-7B-alpaca-vmware
Text Generation
•
Updated
Jun 8, 2023
•
6
ruwan/open-llama-sharded-3GB-7B-alpaca-vmware
Text Generation
•
Updated
Jun 8, 2023
•
5
datasets
0
None public yet