I have found a potentially valid attention mechanism that utilizes the leverage that Aleph Addressing presents, and I have yet to accurately capture until now. I assert, that this addressing and anchoring, presents a hugely increased amount of opinions beyond the current spectrum of explored opinions with shared heads.
https://huggingface.co/AbstractPhil/aleph-splat-0
I dub this attention, Splat Attention. Rorschach splatted differentiated experts, each codebook stacked adjacently, formatted with a methodology of sliding window attention for long context formats. They exist to each have their own view of the same take, and to weigh their views into the outcomes. This allows thousands of codebooks to represent opinions, rather than just one or a few.
Will it work? Yes. Will it work as well as MHA? Probably not. Not yet anyway.
It will at least allow KV caching to exist, while simultaneously enable greater than 8192 depth through a format of utilizable aleph routed attention, specifically curated to align differentiated viewpoints from different codebooks, all represented along the same axis of utilizable space.
The attention mechanism will potentially allow differentiated spaces of 2048 aleph attention heads stacked in unilateral simultaneous execution, which operate exceedingly effectively in preliminary tests - defeating token recall on the tested local and mha heads by a large amount.
What I cannot promise or even say will work yet with certainty, is the depth. The preliminary formulas are converging cleanly and the recall is absolutely fantastic, but that does not mean they will converge yet. This may be the attention that replaces the MHA in these models, and hopefully tonight will have the answer.