ARB_Superbuffers - all the gory details

There seems to be some confusion regarding VBO and superbuffers. The way I understand it, they address two different parts of the GL.

Buffer objects address “byte memory” – GL memory whose fundamental unit is the byte, and it’s the type of memory that we can work with directly in terms of arrays. As with vertex arrays, we can directly access the memory contents of buffer objects and perform selective updates (e.g. cherry-picking just one or two vertices whose attributes we want to change). Traditionally, byte memory objects such as vertex arrays have been stored on the GL client side, usually on the host processor. Sending this memory to the GPU may require extra copying by the driver, which is expensive. The purpose of buffer objects is to extend API in a way such that such memory can be implemented on the GL server side, which on today’s architectures may be AGP memory. You can still cherry-pick individual vertices (in the specific case of vertex buffer objects) and get access to the individual bytes of the buffer object. Until you specify how you want to treat that buffer object (e.g. positions, texture coordinates, normals, etc.), it’s still fundamentally just a collection of bytes.

Superbuffers, in contrast, address “element memory,” whose fundamental unit is some sort of “element.” A common example of this type of GL memory today is the texture object. We have a block of memory that we first feed to GL using glTexImage*, but in that command we specify not only what the data is, but how to format it (RGBA, BGR, ALPHA, etc.). Once the memory is formatted, it forever stays that way. This type of memory is typically stored on the GL server side (e.g. GPU texture memory), and unlike byte memory, you can no longer gain access to the byte representation of that memory once you’ve created it. The superbuffers extension, as I understand it, is a generalization of element memory (previously available in the form of textures, off-screen pbuffers, etc.) so that we have greater flexibility in how we draw to this memory, re-use such memory for different operations, etc. The design of the API is such that (hopefully) the driver can make performance optimizations (e.g. avoiding expensive GPU buffer copies, etc.).

Hope this clarifies some issues. In summary, buffer objects and superbuffers handle different types of GL memory and the ways they are used.

Eric

If that’s true, it sucks. It greatly hinders transition from system memory vertex arrays and VBOs (apps will need ‘regular’ arrays for fallback, VBOs are much easier to implement there; interleaving array data was a good thing last time I checked).

I could have sworn I saw a stride. Maybe I was just looking at ‘size’, thinking it was the stride.

In any case, you’re right; this sucks.

This is, yet more, evidence of the ARB being used as an arena for people to screw over their compeditors. Observe:

If you can’t render to an interleved vertex array, the only way to do render-to-vertex-array stuff is to render to multiple buffers at once (unless you want to run your data multiple times for each component). Of course, ATi is right on top of this, with ATI_draw_buffers. nVidia isn’t. nVidia gets screwed yet again. 3D Labs technically gets screwed (unless they have such an extension), but ATi probably was caching-in political coin for their support of glslang.

Looks like that article was right on the money; the ARB has become a battleground for getting ahead in the marketplace. And, of course, our needs are ignored.

It may be ‘surprising’ for apps when, say, id 1000 joins the pool of defined VBOs as soon as id 1 is created.

To me, that’s a flaw in the whole glGen* structure. glGen* should always be required; they should never have allowed for objects to be implicitly generated simply by using a new integer.

The best way to make these memory blocks work like VBO’s would have been to have functionality for attaching a memory buffer (a leaf) to a VBO object. That way, all your vertex array stuff looks the same.

Originally posted by Korval:
[b]This is, yet more, evidence of the ARB being used as an arena for people to screw over their compeditors. Observe:

If you can’t render to an interleved vertex array, the only way to do render-to-vertex-array stuff is to render to multiple buffers at once (unless you want to run your data multiple times for each component). Of course, ATi is right on top of this, with ATI_draw_buffers. nVidia isn’t. nVidia gets screwed yet again. 3D Labs technically gets screwed (unless they have such an extension), but ATi probably was caching-in political coin for their support of glslang.

Looks like that article was right on the money; the ARB has become a battleground for getting ahead in the marketplace. And, of course, our needs are ignored.[/b]

I think this sounds a little too paranoid. Ever considered that there might be technical reasons behind certain design decisions? It sounds more likely so to me. I think most if not all ARB participants honestly care deeply for the API and works for its best, but of course everyone sees things from their particular vendor’s point of view. And of course everyone wants extensions to work on their particular hardware, even if that may reduce the beauty of the interface. Not all ARB work turns out as perfect, but then there’s no vendor’s hardware that’s perfect and no driver team has infinite resources. Considering the amount of needs, wants and opinions that goes into the work of each extension I think the ARB has actually done incredibly well in most cases.

Yeah.
Also, if I’m not mistaken, the superbuffers thing is preliminary workgroup material and hasn’t been voted on. I see an ATI logo and an ATI employee’s name on the front page, so quite obviously, this is based on ATI’s vision.
All of that may change yet.

I think this sounds a little too paranoid. Ever considered that there might be technical reasons behind certain design decisions?

Certain design decisions, sure. This one, in particular? No.

The hardware certainly doesn’t care whether the pointer is in video or AGP memory, or whether the memory was a render target a few cycles ago; it still supports a stride regardless of the memory’s location. Since the hardware supports it just fine, the API should support it.

The conclusion about ATi’s politicking comes from this fact, plus the fact that ATi’s hardware is not terribly affected by this. Because they permit writing to multiple buffers, one can simply allocate several buffers and write individual components to them. Therefore, the lack of a stride is not a concern for developing on ATi hardware.

This is clearly a bad idea. Yet, here it is.

ATi has done some good things, both for hardware and OpenGL development. However, the fact remains that the ARB has been making some questionable decisions of late, and there must be an explaination for it. Most of these questionable decisions have been falling in ATi’s favor, or at the very least, hurting nVidia.

Do you have a better explaination?

I see an ATI logo and an ATI employee’s name on the front page, so quite obviously, this is based on ATI’s vision.

I wouldn’t say that it is only based on ATi’s vision. ATi is certainly part of the superbuffers workgroup. This document could simply be a demonstration of how far the workgroup has gone. While this could simply be what ATi wants rather than where the workgroup is going, it hurts ATi if this document is vastly different from what the workgroup does. After all, this document has ATi’s name on it; if it’s incorrect, then ATi looks like they’re spreading misinformation.

In any case, you are correct that this document certainly isn’t a finalized spec. And adding a stride is no more difficult than a 2-line change to that spec, so the final spec could very well have a stride. But the fact that the stride is missing is worrysome; how do you overlook such a fundamental part of any vertex array API, unless you specifically removed it for some reason?

Well, you can only write four floats to a render target at most. Since most vertex attributes are four component vectors the limitation about interleved arrays sort of makes sense. Still, it’s weird that you can’t pack i.e. a two component position and u.v texcorods into one four component vector. Having some way of unpacking the data when it’s read from the membuffer would definitely be useful. Of course, both nvidia and ATI support 16-bit float textures, so why we cant use that to pack more data in I don’t know.

Well, you can only write four floats to a render target at most. Since most vertex attributes are four component vectors the limitation about interleved arrays sort of makes sense.

It makes sense to the extent that interleving would not be very useful. It should still be expose, for the day when fragment programs can output multiple components to a single buffer.

Of course, both nvidia and ATI support 16-bit float textures, so why we cant use that to pack more data in I don’t know.

The limitation isn’t one of 128-bits of data; it’s that you can only write out 4 components.

well, not that is might be the fastest way, but since you know the x,y position of the pixel youre writing to, you can easily calculate a vertex in one, and a normal in the next pixel ( odd/even) with an simple ‘if’

Another way to do it is to simply manipulate your data however you want in vertex shaders. Techinically, one could simply pack two 2d texcoords (for example) into each pixel of a buffer and extract them in the shader. Heck, you could even split data amongst multiple arrays and combine them in the vertex shader.

wich is, i guess, the way we’ll have to do it…

and the way i’m doing it even now so i don’t have problems not having a stride…

I’m curious why they have abandoned the previous method of binding GLmem to array, as shown in the GDC03 presentation . After adoption of that older method into the language of the newer proposal, we would get:

glBindBufferARB(GL_ARRAY_BUFFER_ARB, theVBO);
glAttachMem(GL_ARRAY_BUFFER_ARB, theGLmem, whatever);
glVertexPointer(4, GL_FLOAT, your_stride_if_you_insist, offset);
IMO the original binding method is better then newer one, because:

  1. it seems to be more consequent with the rest of SB: you attach the GLmem thing to some existing, currently bound GL object (texture, FB, VB), rather than directly to the GL_VERTEX/COLOR/etc_ARRAY targets. (As far as I understand the spirit of the proposal, GLmem is not intended as a new kind of GL object, but as a holder of contents of other objects - otherwise I don’t see justification for the new GLmem type, separate namespace, and lack of glGen)

  2. it’s easier to use: only initialization code is changed with SB (glAttachMem vs. glBufferData), but later you use all VBOs in exactly the same way.

  3. there is no need for new gl(Multi)DrawElements() functions

  4. the stride problem disappears: you set the stride as in regular VBO (although I’m not terribly convinced about the need of the stride in render-to-vertex-array usage)

So I’m curious, what adventages of the new binding method (glVertexArrayMem) outweighted these 4 above, as I simply can’t imagine any…

[This message has been edited by MZ (edited 08-20-2003).]

Is it possible that the final specs will get changed or will OpenGL begin to suck as DX does a long time ? I can’t understand why every ARB_Extension needs another way of handling the whole problem. Why not using the VBO and SB the same way ? More and more i think, that DX will be the better choice for the future.
I begin to see it this way :
ARB is too slow to handle the GPUs evolution.
+
The time for Vendor-specific Extensions is over (too much work, i can have the same in DX with only one renderpath).

Am I the onlyone or are more people angry about the way OGL gets killed (hope MS won’t be so successful). OpenGL gets too complicated. No more rapid-dev is possible and that was why i used OpenGL.

Please say that the specs are not ready. PLEASE !

cu
Tom

Am I the onlyone or are more people angry about the way OGL gets killed (hope MS won’t be so successful). OpenGL gets too complicated. No more rapid-dev is possible and that was why i used OpenGL.

The super-buffers extension, in and of itself, is quite simple. As for OpenGL getting too complicated in general… tough. Complications are needed to be able to access advanced features.

Please say that the specs are not ready. PLEASE !

If you weren’t paying attention, these aren’t specs; they’re a powerpoint presentation that ATi made at Siggraph. All we know is that it represents a particular state of the super-buffers extension.

BTW, here’s the most likely reason why they don’t use VBOs.

In general, it’s a bad idea for certain functionality in an extension to rely heavily on an unrelated extension. For example, you wouldn’t want render-to-vertex-array to rely on an extension that you can’t be sure is implemented everywhere.

Note that VBO was only selected for insertion into the core a few months ago, and that vote was only preliminary. Back then, the superbuffers workgroup couldn’t count on VBO becoming part of the core. So, they had to build into their API a vertex binding method.

Now that VBO has been (preliminarily) approved for insertion into the core, it is now reasonable for superbuffers to rely on that functionality as an interface for rendering meshes.

Originally posted by MZ:
(quote from proposal)– Memory objects attached to a framebuffer object must have the same dimensions and dimensionality

DX9 is less restrictive: (from SetRenderTarget() description) “The size of depth stencil surface must be greater than or equal to the size of the render target”. Since HW supports it (wouldn’t be in DX otherwise), then why not have this exposed in GL too?

excuse me if I reply to my own post in month-old thread, but: I’ve just spotted evidence , it looks like the technique is nothing unusual in DirectX.

[This message has been edited by MZ (edited 09-02-2003).]

So what’s wrong with viewport and scissor? What is the difference here in the context of what this guy is doing? Doing this sort of thing and trying to pretty it up behind an abstract interface that merely looks flexible can often lead to hidden performance pitfalls.

So what’s wrong with viewport and scissor?

It is often the case that, in our quest for flexibility, we forget the capabilities that we already have.

Granted, viewport and scissor don’t actually prevent us from allocating a 1600x1200 colorbuffer, to use with a 1600x1200 z-buffer, but it does effectively let us section one into smaller pieces.

You can’t allocate two 512x512 buffers, one as a color buffer and one as an aux buffer, and render to both (with ATI_draw_buffers, for example) with a 1600x1200 z-buffer. You’d have to allocate two 1600x1200 regions and section them with viewport and scissor. And it has to be 2 because you can’t have two buffers that draw to different regions.

There’s still usefulness in having the flexibility.

I don’t understand what you are talking about, dorbie, viewport and scissor are completely unrelated to the problem.

To generalize what the guy from my link is doing, consider this: You have a number of render targets. They can have different sizes and different dimensionality (RECT, 2D, CUBE). When rendering from any them, you use only color data (let’s forget about ARB_shadow). When rendering to any them, you temporarily need depth buffer (+ stencil &| multisample). You can choose one of two strategies:

A. Create one depth buffer for every possible color buffer size (and every possible mipmap size). A group of equally sized color buffers can share one depth buffer.

B. Create single depth buffer of size being maximum of all possible color buffers, and use it when rendering to any of them (that’s what the guy in the link is doing B)

With current GL neither of above is possible (except very limited A, with double buffered WGL_RTT trick).
With DirectX both strategies are available.
With current SB proposal you can do only A, B is explicitly disallowed - and this is my concern.

You can’t mimic the B strategy with scissor & viewport, if this is what you’ve been trying to suggest (unless you considered glCopyTexImage as RTT method, but this is not related to SB). And it is not all about “trying to pretty something up”, but about exposing a HW feature (ability to render to differently sized color and depth buffers), which appears to be widely available (it’s not even optional in DX, it is standard), and which can be useful.

I thought about possibility of solving the problem automatically inside driver - there is “undocumented” glInvalidateMem(GLmem mem) thing in the proposal. With well placed hints saying “From now I don’t care for contents of this depth buffer” the app could pretend it is doing A, but the driver would turn it to B internally - something like vertex buffer renaming when “discard” hint is given by app. With this, the extra flexiblity of DX wouldn’t be necessery to achieve the same effect, so maybe the ARB guys have already gone this way.