glDrawElement with independant vertex, normal and texCoord indexes

Originally posted by Korval:
DHEUDJKE to you too.

As I said, I Am Not A HardWare Designer (IANAHWD)…

Sorry for being terse earlier (ahem), my point before was that 4-bone skinning would barely fit using the aforementioned hypothesized multi-index approach (i.e., where the matrices would get sucked into the VP, masquerading as attribs).

I think we agree using up attribs this way is not ideal, but I’m not convinced we’re going to see arbitrary bindable AGP that’s as fast constant registers or vertex attribs bound via multi-index.

Maybe an actual HW designer can speak to this, but fast VP execution would seem to require not waiting on any AGP memory fetches. I don’t know. Does the VP itself do the job of pulling the AGP vertex data? I doubt it. I’d bet it’s an earlier instruction block and the VP silicon is tuned to work entirely in register space. (the actual answer is probably proprietary, so I don’t really expect much, but here’s to asking).

Anyway, if there’s a reasonably large cache to ease the pain of AGP reads, fine, but that’s still a fixed window–matrices might be very spread out in AGP, meaning a potential for multiple cache fills per vertex program cycle. Sounds painful.

However, if the memory gets pulled into the VP via multi-index, that pre-fetch can happen before the VP even starts. The pre-fetcher block can be doing nothing but look at indices, read AGP, and assemble vertices into a transform queue for VP or standard T&L. Seems much more deterministic and much simpler to my naive eye. And it sounds like what would be happening anyway with a single index reading from arbitarily placed AGP (non-interleaved) arrays.

Avi www.realityprime.com

Hello Everybody

I personnally think that we are missing the point. This should be about the feature and not the current design of the card.

I personnally think that multi indexes adds a lot of flexibility to my model loading and it could possible save bandwidth going to the card.

After that its the manufactures responsibility to make it fast. NVidia and ATI have some smart engineers working there. Leave it the them to figure out the best way to optimize.

Lets discuss the idea and possible good point and bad points of why multi-indexing is a bad idea. Once it goes to the review board you can be damn sure that hardware companies are going to be regecting this is they don’t have a way to optimize it.

Ben

Originally posted by zander76:
[b]Hello Everybody

I personnally think that we are missing the point. This should be about the feature and not the current design of the card.

I personnally think that multi indexes adds a lot of flexibility to my model loading and it could possible save bandwidth going to the card.

After that its the manufactures responsibility to make it fast. NVidia and ATI have some smart engineers working there. Leave it the them to figure out the best way to optimize.

Lets discuss the idea and possible good point and bad points of why multi-indexing is a bad idea. Once it goes to the review board you can be damn sure that hardware companies are going to be regecting this is they don’t have a way to optimize it.

Ben[/b]

I don’t know. Given there are alternate ways to do the same thing, it may come down to what gets the best results for the least effort, and that involves weighing API changes and cost to implement in HW. I don’t see a lot of good winning a new extension that’s not supported because it’s too expensive. So yes, NVidia and ATI can undoubtedly do better than us, but it’s also important to convince people this is feasible and worthwhile.

One thing that did occur to me the other day (possibly to other people too) was that the existing single-index cache structure might be sufficient for caching if there’s an easy/fast way to create a unique hash of multiple indices into one. Seems like that would do the trick, however uniqueness is a non-trivial problem. Just an idea.

Anyway, I think we’ve given some good agruments as to why multiple-indices don’t always save bandwidth, so you may not want to keep falling back to that without your own examples. In fact, without AGP’d indices, it would probably be more expensive to separately index texcoords (8 bytes) and colors (4 bytes), for example.

But it could be a big win, IMO, for bigger shared data, like matrices, perhaps even normals. The biggest win, IMO, is in being able to pull vertex-indexed data into vertex programs without wasting massive per-vertex storage or stopping/starting draw batches to reload VP constants.

So unless you want to make some new arguments or someone else wants to argue the previous ones, I’m not sure what else we can add at this point.

Avi

[This message has been edited by Cyranose (edited 08-26-2003).]

The way I see the multi-index feature is that we will end up in using an index table to lookup in all the index tables, thus index of index. That’s not likely to be an optimization at all, in my humble opinion.

If this can no to be in the GL core for performance reasons, this can perhaps to be included into one standard OpenGL library add-on like GLU or GLUT ?

Hum, and why NVIDIA prefer to use a cubemap texture for vector normalisation if the table lookup path if as bad as this on modern GPU ???

And somes SSE/3DNow! function don’t work internally with some table look-up ???

And what about the use of a palette with CGA/MCGA/VGA cards if the index lookup is as bad as this on old/current/next hardware ???

And for me, 4 indices of one byte each (cf. a 32 bit value) can certainly be transfered more speed that (4 float for the position + 3 float for the normal + 4 byte for the color + minimum 2 byte for the texture coords), with or without AGP …

@+
Cyclone

“Infact the biggest problem is bus bandwidth and not the processor on the card.”

“No, the biggest problem is with the post-T&L cache”

And why not cancel this [in]famous post T&L cache if the GPU is now more speed that the memory ???

In fact, the problem is REALLY in the input of the T&L stage (cf. load from memory/disk to GPU memory), it is not at the output (cf. the final fragment’s values computed on the GPU and used in the rasterisation stage)

@+
Cyclone

“The way I see the multi-index feature is that we will end up in using an index table to lookup in all the index tables, thus index of index. That’s not likely to be an optimization at all, in my humble opinion”

With the language C/C++ this type of instruction is really used frequently :

   **a = something

or

    something = **a

And if it is really frequently used, it isn’t for nothing …

But don’t have the texture stage something that is already equivalent, cf. dependent texture fetch or something like that ?

@+
Cyclone

Originally posted by vincoof:
The way I see the multi-index feature is that we will end up in using an index table to lookup in all the index tables, thus index of index. That’s not likely to be an optimization at all, in my humble opinion.

That seems like a good idea to me. I can see the initial lookup being more expensive to resolve because of the extra indirection, but if you implement it as you said, the post T&L cache cost should be the same as now and the cache-hit ratio should be at least as good.

The bigger question I have is API. Ideally, the extra indices (I’ll call them indirections so as not to confuse with the master index) would need to be flexible. For example, I might want three attribs to be tied to one indirection but a fourth to have it’s own indirection. Plus, we’d want the indirections to live in fast memory. So maybe something like an enable/disable pair plus normal glDraw* calls plus the equivalent of

glBindAttribIndirect(attribute_number,indirection_table)

…where the programmer could use the same table for two or more attributes if she so chose. I also suppose it would be possible (though perhaps not useful) to default to indirection = master-index for any attribute where a table isn’t bound.

Avi
www.realityprime.com

[This message has been edited by Cyranose (edited 09-05-2003).]

glBindAttribIndirect(attribute_number,indirection_table)

where attribute_number can be something such as GL_POSITION, GL_NORMAL, GL_COLOR or GL_PARAM0_EXT, GL_PARAM1_EXT, …

And when indirection_table is set to NULL,
a default table of incremented indices from 0 to 65536 is used

Yes, it’s something like this

I still don’t see immediate use for this, but

Originally posted by cyclone:
And when indirection_table is set to NULL,
a default table of incremented indices from 0 to 65536 is used

Nope. You’d use GL_NONE (which is zero, too) and make that behave like pass-through. Some people use more than ushorts for indices