Rayen
Lissanro
AI & ML interests
None yet
Recent Activity
upvoted a collection about 1 month ago
GigaChat 3.5 liked a model 2 months ago
prefeitura-rio/Rio-3.5-Open-397BOrganizations
None yet
Nice mockery from QWEN team they release the 2.4T model nobody can use but the 27B is 2 days delayed. WTF
👍 2
4
#1 opened about 19 hours ago
by
akierum
awq 4 bit
3
#1 opened 4 months ago
by
hamadfx
Please release a model with native 4-bit quantization
👀➕ 11
4
#4 opened 6 months ago
by
calycekr
128K context does not work (possibly because YaRN meta information is missing?)
➕ 1
2
#8 opened 10 months ago
by
Lissanro
MXFP4_MOE
🔥🚀 1
11
#1 opened 12 months ago
by
marcelone
Incorrect Model Uploaded
🤗👍 18
6
#8 opened 12 months ago
by
noteventhrice
Context length: is it 128K (as mentioned in the model card) or 160K (as specified in config.json)?
1
#17 opened 12 months ago
by
Lissanro
Any plans to release an updated version based on DeepSeek-V3-0526 + R1, or how to create the merge myself?
14
#4 opened about 1 year ago
by
Lissanro
Please consider creating ik_llama.cpp compatible quants (without llama.cpp-specific MLA tensors)
1
#1 opened over 1 year ago
by
Lissanro
chat_template.json is missing
2
#1 opened over 1 year ago
by
Lissanro
chat_template.json is missing
2
#1 opened over 1 year ago
by
Lissanro
Tell me how do you feel about this model without telling me how do you feel about this model
4
#5 opened over 1 year ago
by
MrDevolver
Is this model native 128K context length, or YaRN extended?
7
#28 opened over 1 year ago
by
danielhanchen
Doesn't Generate `<think>` tags
3
#25 opened over 1 year ago
by
bingw5
Works great on oobabooga, but always ends with assistant
1
#3 opened over 2 years ago
by
Noodlz
This could likely be dewokefied and possible even improved using mergekit's new 'Model Stock' method!
🔥 2
31
#5 opened over 2 years ago
by
jukofyork