interesting. I mean day-to-day usage is a requirement if you want to buy some of these machines for real. But good to know that graphics works now. Last time I checked it wasn't and that was significantly after launch
The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one
Yep more hops from the lower Q is likely going to skew the vectors further over time.
I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?
The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.
This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.
I wonder how well it performs in practice, because I can't help seriously doubting those benchmarks. That would put this 6GB model in Opus 4.6+ ballpark. Granted, that's mostly to Qwen 3.8's credit, but it's hard to believe that Qwen's already unbelievable capability density can still be compressed this much more.
Genocide is usually not an efficient way to wage war. Even ancient armies usually did not slaughter entire towns, not because of ethics, but because fighting sucks and if you slaughter everybody you come across, nobody is going to surrender and every person you come across will fight you to the death.
I say usually, because there are some exceptions and nukes are one of them, they make genocide a viable war strategy. But even nuking Vietnam would not have worked, cause you will simply get nuked back.
In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.
At least, as an ignorant outsider, that's how it seems to me.
As I understand it, that setting does not prevent them training on user data, just which derivatives are used (i.e. just PII scrubbed vs certain types of synthetic summarization)
That is an absurd and entirely untenable position that breaks with approximately all western conventions.
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
I agree it's an issue but maybe we are going at it from the wrong angle. Why just be scared of the Chinese? We should protect ourselves from both domestic & foreign adversaries.
Any electronic device with vision, sound, or radio communications capability should be required to be source-available or at least independently audited. Proprietary code must run in sandboxed black-boxes that can't directly communicate/interact with the outside world.
Sounds like an impossible ask? Well, banning Chinese robots/routers is also an impossible ask.
That's their point though. Basically no modern phone/laptop/tablet other than Apple offloads audio decoding (of any codec) to hardware. You can check this on Android phones by installing the Codec Info app.
I’m pretty sure no x86 chip has hardware decode/encode for audio. Together with dGPUs, they tend to have decoders for JPEG and decoders/encoders for H.264, H.265, AV1 and sometimes VP9.
There is absolutely a crisis of bad health outcomes due to relatively simple home medical procedures. E.g. People take Tylenol at home thinking it is safe, meanwhile Tylenol posionsing is the second most common reason for liver transplantation worldwide.
The reason we accept these crisis is due to societal, cultural and religious tradition/pressure. IMO, in an ideal world, many of these things should draw additional scrutiny.
FUTO Swipe supports ClearFlow, which is exactly what you are talking about, a keyboard layout optimized for swiping: https://clearflowkeyboard.github.io/
I switched to ClearFlow a month or two ago after learning of it on Hackernews. It is available in GBoard.
I'm happy with the switch. Like any keyboard switch (I've gone from Qwerty to Dvorak and now a Colemak-dh derivative with about ten years on each) it takes some time to learn the layout. Overall I'm happy with it though and there are less frustrating misinterpretations and corrections needed.
This post was swiped on it with only two corrections and the second one was my fault as i misremembered a key location.
There's a section [1] on the page that has instructions, and video [2] too. I had to select the English (US) language to get the option to select ClearFlow.
Thanks. It's available only for the US layout, not UK.
I'm writing down a few impressions:
- the layout is unusual, but I get the motivation. Distances are minimised and letters are arranged so that ambiguity is removed.
- although I'm very slow, I haven't made a single mistake so far. Clearflow allows me to swipe much more accurately than stock gboard.
- the square keyboard layout unfortunately means that half the letters are constantly hidden behind my thumb. As I'm unfamiliar with the layout, this means that before swiping a word, I have to look at the layout, memorise letter locations and plan the movement
- since I write in multiple languages and Clearflow is available in only one of them, I would have to memorise a completely new layout for a language I write in only half the time.
Hi,
Yes I'm in the same boat as you - had to switch to US language instead of UK. I've been addiing the anglisised versions of words to my dictionary as I go along so it's becoming less of an issue over time. Maybe I'll switch to FUTO in order to not have to deal with this anymore. Gboard has one nice feature though in that I have multiple languages enabled so I get correct predictive completion in non-English languages.
For learning ClearFlow, I used the Games app available from the "Clearflow Games" section on their website: https://clearflowkeyboard.github.io/
I also have the issue of the thumb getting in the way so I spent a couple of days playing the games to get my layout memory up and then it became usable without frustration and I'm not looking back now although I occasionally still forget the odd letter location.
I have been a ClearFlow user for over a year now. Generally I like it, but there absolutely are still common words that are hard to input consistently. The THEA cluster has given me no shortage of problems. Still a fan though.
I actually can't find the answer on either of the linked pages, so it would be good to know. And I think people's experience is more important than the claims in these discussions.
And anyway, there's no keyboard on earth who can handle multi-language typing in a sane way. They either mash all languages together, or force random layouts on you, or... I stuck with GBoard because I just hate it less than others, so when I found this topic I thought yay let's try - until I read it's only for English. So there.
I really like how gboard handles it. It figures out what language I'm typing from the first couple of words and prioritises those. This way I can even mix languages within the same sentence and it will still recommend the right ones. It's really really good.
Are you sure, like really sure you mean gboard? In my gboard I choose the language I want to type in, and that's the language for everything (settings, typing, recommendations...).
I'd like to know where did they get the stats ClearFlow mentions in their site (reducing backspace corrections by 37.5% and shortening finger gliding distance by 41.6%.) and see what method did they use to analyze those swipe patterns and create .
It could be interesting for applying it to different languages (or modified word corpus).
The thumb typing muscle memory does not translate to finger typing at all. Most Dvorak or Colemak users are comfortable using QWERTY on their phones. Clearflow really only works well with swipe.
No, not muscle memory, but at least the idea of knowing where keys are. I'd bet that non-qwerty typers mostly started with qwerty and possibly still need to use it on some occasions, so they remember.
I can't speak for everyone but I did an experiment to see if this was the case (I actually just wanted to switch to Colemak everywhere. But iOS doesn't support it natively and third party keyboard landscape is pretty bleak on iOS). I switched to Dvorak on my phone and got really comfortable with it. I was already comfortable with Colemak on my computer. Then I tried
1. Switching to Dvorak on my computer AND
2. Switching to Colemak on my phone
I already had a feeling that my Dvorak experience on phone wasn't gonna help typing in a physical keyboard and indeed that was the case. But also, the physical keyboard experience with Colemak didn't really help with thumb typing/swiping on the phone. On a physical keyboard, I can only type QWERTY at around 30WPM and Colemak at around 130WPM. However, on the phone QWERTY is my strongest layout (even when I can only type at around 50WPM with it).
Then there's the design philosophy. Colemak aggressively keeps the most used letters on the homerow. Paste a random Wikipedia article into this page and you'll see (https://www.patrick-wied.at/projects/heatmap-keyboard/ alternatively, here's the heatmap for the introduction to today's featured article about Augustus: https://postimg.cc/9RT8P4N6). On a phone this means that a lot of the times you'll be swiping back and forth horizontally. Resulting in identical swipe paths for a lot of common words.
And if Clearflow was ported to a physical keyboard, most of your fingers won't be doing any work and the active fingers will be doing weird tap dance. The position of spacebar (undoubtedly the most used key on the keyboard) to the sides where you'd be using your pinky would make it an RSI machine.
My point is that if you're comfortable with QWERTY on the phone, that's not because you're comfortable with QWERTY on the desktop. That's because you got used to QWERTY on the phone.
> For example preventing people who have repeatedly been convicted of violence ("hooliganism") from entering sports stadiums.
You should't need face recognition to do this (as it seems stadiums are already successful in banning people who haven't paid from entering the venue).
But it can be annoying to use day-to-day since a lot of stuff is not upstreamed or not integrated into the major distros.
reply