Originally posted by Korval:
DHEUDJKE to you too.
As I said, I Am Not A HardWare Designer (IANAHWD)…
Sorry for being terse earlier (ahem), my point before was that 4-bone skinning would barely fit using the aforementioned hypothesized multi-index approach (i.e., where the matrices would get sucked into the VP, masquerading as attribs).
I think we agree using up attribs this way is not ideal, but I’m not convinced we’re going to see arbitrary bindable AGP that’s as fast constant registers or vertex attribs bound via multi-index.
Maybe an actual HW designer can speak to this, but fast VP execution would seem to require not waiting on any AGP memory fetches. I don’t know. Does the VP itself do the job of pulling the AGP vertex data? I doubt it. I’d bet it’s an earlier instruction block and the VP silicon is tuned to work entirely in register space. (the actual answer is probably proprietary, so I don’t really expect much, but here’s to asking).
Anyway, if there’s a reasonably large cache to ease the pain of AGP reads, fine, but that’s still a fixed window–matrices might be very spread out in AGP, meaning a potential for multiple cache fills per vertex program cycle. Sounds painful.
However, if the memory gets pulled into the VP via multi-index, that pre-fetch can happen before the VP even starts. The pre-fetcher block can be doing nothing but look at indices, read AGP, and assemble vertices into a transform queue for VP or standard T&L. Seems much more deterministic and much simpler to my naive eye. And it sounds like what would be happening anyway with a single index reading from arbitarily placed AGP (non-interleaved) arrays.
