unique ROAM VBO issues and a clincher

You know, maybe we should get off Michael’s case a bit.

Michael comes across as an “architecture astronaut” - head in the clouds and having difficulty explaining himself in terms understandable to us mere mortals.
I only get about half of what he’s saying - some pictures would probably help a lot.

But that doesn’t necessarily mean his idea is useless. For example, what use was there for fractals until someone thought of using them for compression?

Many people are looking at this assuming that the CPU would have to do a lot of work.

Suppose that precalculating everything in this way IS a smarter approach than just trying to render every triangle, or any of the other LOD mechanisms.

Also suppose that this is a cheap method to select those precalculated tris.

And suppose it is simple (and cheap in terms of space) to implement in hardware.

On those conditions it might well be interesting

Of course it requires Michael to work out the details, and it may well not be worth it on any current or currently planned hardware, but that’s not necessarily the point of research.

Michael, it may be an understatement to say Knackered doesn’t play well with others.
Usually I just skip his posts, but he (well, I assume it’s a he) did give you a couple of useful suggestions.

Originally posted by michagl:
oh, and i read the nvidia architecture paper… i was surprised how well my preconceptions of the hardware aligned with the info there… pretty much spot on actually.
Now I don’t know about the rest of you, but I for one am really impressed by this.

Originally posted by T101:
Michael comes across as an “architecture astronaut” - head in the clouds and having difficulty explaining himself in terms understandable to us mere mortals.
I only get about half of what he’s saying - some pictures would probably help a lot.

No, Michael comes across as someone who thinks everyone is stupid, except him.

I think he’s Derek Smart in disguise.

/A.B.

there is a lot to respond to since my last post.

Originally posted by knackered:
[quote]Originally posted by T101:
Michael comes across as an “architecture astronaut” - head in the clouds and having difficulty explaining himself in terms understandable to us mere mortals.
I only get about half of what he’s saying - some pictures would probably help a lot.

No, Michael comes across as someone who thinks everyone is stupid, except him.
[/QUOTE]quickly, i don’t believe ‘everyone is stupid’… i’m here specificly because i assume at least most of the people here know this stuff better than i do as far as the hardware is concerned. and in my defence, i do my best to keep my distance from people and contemporary television programming, so my socialization might seem strange to you. just think of me as being from another culture than yours. take my words very literally. i’m not a silicon valley lounge lizard.

my cpu right now i believe is around 2.8Ghz, 2.5 at the minimimum – AMD.

i’m sorry i don’t have the time and energy to describe this stuff better… but i also very strongly get the impression that possibly interested parties are really not trying. i’m willing to meet people half way, but i really don’t think this is an appropriate forum for an indepth explantion. and frankly my explanation has been at least as explanative as most acedemic papers, which generally do very little more than to scratch the surface.

after i feel like i’ve followed this line of thought as far as it can go, and have a release build to demo, then and only then will i look into collaborating with people to help really tweak it (and keep it tweaked) for hardware, and help develope rigorous arguments on paper. but i develope a lot of systems, i have at least 10 very exciting projects in development right now… ranging from revolutionary operating system environment design (designed to run atop windows, and debian linux so far, though indapendant operation is the goal ), to an equally exciting GUI system, an dynamic universal interterpreter capable of parsing any unambiguous sequential language(s) interoperably (currently does basic c++, lisp, and opengl), a sleu of graphics systems, a 2D dynamic disk system for streaming massive maps to disk with an opengl interface, which basicly treats one map as the frame buffer, and the other as a texture, with a bios system inbetween… bios targets can be ram to disk, disk to disk, or disk to ram. and other much more exciting work which i’m not at the liberty to discuss. after Genesis is mature enough, i’m planning to set my targets on what will essentially be a portaling system called Daedalus… i won’t give myself much time to concern myself specificly with papers and what have you… i’d rather be building something tangible. its a lot more complicated than just this naturally, but for what its worth everything is scheduled for free release.

as for new hardware running 10x faster… please give me the price points on those cards.

i guess my biggest issue however with this many triangles, is i don’t see how disk io can keep up with these numbers. it seems like heavy load times would be in order, especially for procedurally generated data. maybe Nintendo style RAM cards will go back in style… because from what i hear, disk io is not keeping up.

finally, i’m trying to cope best i can with these numbers… but even if the algorithm can support infinitely high batches. it seems like the bottleneck would be in getting that data into ram in the first place.

and yes, i do believe the algorithm could be mapped completely to chip. if anyone at nvidia wants to talk, i will try to draw up a more envolved explanation, about how i would imagine the chip. it would make an awesome hardware nurbs solution if nothing else. also, i should add, that i do not believe there is a computationly more efficient way to go about ROAM. so if this algorithm fails, so would the future of ROAM i believe. but i don’t believe ROAM will fell, i imagine it will be the cornerstone of graphics come the end of time.

finally, there is a batch solution i intend to impliment next. i plan to replace all the driver dispatches with a single glDrawLists dispatch, and two bytes per mesh in 255 block chunks. each byte will be a display list, the first will transform the modelview into local space if desirable – ie hold a matrix – as well as the basic glDrawElements set up. the matrix might be optional, it helps precision for large models, but for smaller models with possible skinning, world space (no matrix) would be more desirable. the second list per mesh will probably just have a single glDrawElements.

the lists would just be recompiled as necesarry (fairly rarely). list zero in each block would be a NOP list (it does nothing)… meshes not visible to the frustum would be set to zero in the list string.

so in the end, all driver instructions would be reduced to a couple glDrawLists calls, and some short byte strings to be passed through AGP.

sincerely,

michael

before i log off real quick i just wanted to say that most everyone contributing here has been of great service, if only because it got me thinking very criticly. many developments have emerged since the beginning of this thread, which if nothing else would’ve taken some time longer to emerge were i not to have stumbled into this forum looking for a solution which ultimately turned out to be as trivial as vsync.

and yes, i apreciate knackered even i think, to the extent that he has maliciously attempted to corner me, and forced me to adapt my critical thinking… but then on the other hand maybe i’m just confusing him with other contributors, and he really only served to muck up the discussion… but at least maybe he spiced the thread up enough to rally some onlookers.

all said, i apreciate your efforts… i realize the personal ‘costs’ of your contributions… and i’m very grateful. i only hope my personal maneurisms do not offer the wrong impressions.

if any ill will is detected in my words, i’m sure it is only a misunderstanding. i might be criticly honest, but i’m not spiteful.

sincerely,

michael

Originally posted by T101:
some pictures would probably help a lot.
there have been two images shared on my behalf throughout the thread, which do well as diagrams.

i’m assuming if you’ve read from cover to cover you are familiar with them… but just to reiterate, here they are again:

http://arcadia.angeltowns.com/share/genesis-mosaics-lores.jpg

http://arcadia.angeltowns.com/share/genesis-mosaics2-lores.jpg

maybe later i will share some older more picturesqe (photo-realistic) images.

I’ve just realised I haven’t contributed anything positive to this thread in…a good while now, so I’ll leave you alone from now on. This is most certainly my last “contribution” to your thread. Enjoy the peace and quiet so you lot can continue to thrash out the nitty gritty implementation details of this algorithm nobody but michael understands.
I’ll go and tease some beginners before supper.

Originally posted by Adrian:
[quote]Originally posted by michagl:
I’m using an nvidia QuadroFX500.
How does your terrain engine cope with one million polys on screen, because that is what the next generation of terrain algorithms should be achieving. Your engine would need to do half a million draw calls per second for 30fps.
[/QUOTE]i don’t think a good ‘terrain’ needs 1 million polygons to be pixel perfect on forseeable displays. i would guess that most of those triangles would be wasted power. with that many data points, you may as well not even render triangles… a point renderer could probably saturate realistic displays quickly. just throw a bunch of points at the framebuffer, and then run a pass through blending the one or two pixels that didn’t get hit.

my big question is how the hell are you going to get that much data into ram, without making users wait minutes for it to load up. and finally how, long will users put up with very small flat world partitioned environments… and that is just discussing terrain engines. i bet those online rpg players would like a planet they can circumnavigate seamlessly, and whoever really delivers that first, nothing less will suffice from then on.

flat worlders be damned.

the days of asteroids are comming to a close.

and i’m building for a system i won’t have to replace till the end of time. i’m not crazy about tackling the same problem twice.

Michael, I had a chance this morning to play around with Sean’s demo, over at Gamasutra. It’s interesting–flying about a solar system, bouncing into planets–a jolly good time. I admit I haven’t given spherical LOD much thought, as I see quite a few obstacles to overcome in certain game-world scenarios. And you’re right, the spherical LOD section at vterrain is beginning to atrophy (the aforementioned demo is from 2001).

That notwithstanding, I don’t see why some of the more hardware amenable algorithms couldn’t be applied to a sphere, or any other geometry for that matter. Consider, for example, geo-mipmaps being bent onto a sphere. Additionally, I don’t see why an image based system wouldn’t work, provided a good metric could be found. Or clipmaps: Again, if a suitable metric could be provided, who knows. The point is that these are all very hardware friendly approaches to detail reduction.

It’s just that after playing with ROAM again, I’m reminded of the major drawback of its premise, which is to minimize rendering costs at the expense of the CPU. This just doesn’t fly with modern hardware. Unless, of course, you can manage to alter the hardware. Even with a LUT: this would be essentially tantamount to an image based system. The link given earlier describes such a system. If one were to render multiple sections and warp them onto a sphere or cylinder, for instance, this could work quite nicely, and be very (ultra-modern) hardware friendly. In both cases, the warping is uniform, so the side-effects would be marginal, and would depend largely on the scale. Other geometries could be problematic, however, if wapring artifacts are to be minimized. But by and large, its seems to me that the most effective algorithms are simple, and fit very neatly with the way modern hardware works.

I think hardware innovation is a real liberator. The algortihms in use today are only possible with the hardware that supports them. They simply don’t make any sense in a vacuum. By not taking advantage of these advances, I think you do yourself a disservice. Perhaps in the distant future, when most, if not all graphices code is executed on the GPU, then a ROAM-like approach might find its way back into the fray. But my vision of the future is far simpler. I envision increasinlgy simple algorithms executed by increasingly prodigious graphics chips.

Anyway, I’ve enjoyed this dicussion. This is all such entertaining stuff.

Incidentally, the cylindrical world you mentioned reminded of the book “Rendezvous with Rama.” Don’t know why they haven’t made a movie out of that one :slight_smile:

[edit: format]

You may just get your wish:
Rendezvous with Rama

i don’t think a good ‘terrain’ needs 1 million polygons to be pixel perfect on forseeable displays.
I’d say that a million is a bit of a stretch, but 500,000-700,000 is not unreasonable. And 1 million is good for the ambitious developer.

how, long will users put up with very small flat world partitioned environments… and that is just discussing terrain engines.
There have been a number of games in the recent past that stream data rather than having specific loading points and partitioned levels. Dungeon Siege, for example. GTA3+, for another example.

Roam is not required for this. Only a system for streaming static mesh data from the disc is needed.

Originally posted by michagl:
i don’t think a good ‘terrain’ needs 1 million polygons to be pixel perfect on forseeable displays. i would guess that most of those triangles would be wasted power.
Display technology is moving quite fast now. You can buy a 24inch Dell LCD 1920X1200 for $1200 US. I’m sure the price will fall quickly this year like the 20inch version did last year.

It’s better that the gpu at least has the opportunity to waste some vertex processing power rather than sitting idle. Despite the wastage the brute force methods will still render more useful tris/sec that a cpu heavy approach.

The 10x faster card I was referring to was an ATI X850XT platinum capable of transforming 800M vertices/sec, compared to the 40-80M of the FX500. (I’m not familiar with the FX500 I found conflicting specs of its transform speed). I believe the X850XT sells for around $500 US.

Originally posted by knackered:
I’ll go and tease some beginners before supper.
That’s the spirit ol chap! :smiley:

-SirKnight

not to be fececious, but you are drasticly over simplifying the problem graham. its easy to sit back and play devil’s advocate when you have nothing to loose… but the truth is if it was as simple to produce a real ‘game’ under such constraints, it would’ve been done a few years ago. i really don’t have time to get into the logistics of it all… but graphics programmers are really addicted to the plane, and its counter part euclidean geometry.

doing everything on the gpu is fun for tech demos, but highly impracticle for a real simulation environment.

as for streaming… yes i surely hope all exploration based games utilize some form of streaming. but consider the example of 1 million polygons… unless you are going to force your players to travel down long corridors or whatever to get to a ‘save point’… this isn’t going to work out. if the player could save in the field, the system has to be prepared to spit those 1 million polygons in view out immediately… and then grab its ass and do its best to stream the polygons which will soon come into view in asap.

there is no way to accomplish this without some discritized LOD based system… and by the time you have that, you are halfway to a ROAM solution. taking the last step, is just a matter of relying on your brain rather than abusing the gpu. but that doesn’t begin to change the fact that you have to get those 1 million triangles on the screen and fast. hard disks aren’t even that fast in my experience… that effects clod or not. the difference between continuous and non-continuous LOD, which is the only remaining question for terrain, is that non-continuous LOD relies on artificially inflating the triangle count to try to obscure seams. in fact it marvels in the fact tthat its triangles are only a pixel big… because if it was bigger than that, the transition would be obvious.

this kind of thinking appears to be the only way out. granted ROAM algorithms in the face of hardware have not been successful… that is why i’m presening a novel aproach which really walks right up to the effeciencey of static meshing, but comes with much less of a disk IO overhead, and saves geometry for elsewhere. if you understood how it worked, you would see just how non-existant the cpu overhead is. like T101 i believe contributed… everything is precalculated. selecting which ‘mosaics’ to tile your surfaces with is practicly a non-event that can be tucked in just about anywhere, and optimized to non-existance for a slight drop in precision.

finally as for comments about tailoring a system to spherical or cylindrical geometry. first of all this is not nearly as simple as it might seem on the face. secondly, i’m not trying to turn out ‘a’ game or something in the next quarter. i’m building a robust system which can be used to produce any ‘game’ with a minimal code base and maximally abstracted and detangled interface. so for me, a robust solution wins out over a specialized solution. i have no direct plans of trying to turn a profit, so i have no reason to produce disposable code. i’m building a tool which will be ready for next generation hardware when it hits the ground running, while the rest will be still very invested in archaeic highly hardware dependant code bases. hell most code bases are built to be disposable… this is so wasteful, because if people have half the sense they like to think they do, we shouldn’t still be compiling games in 2005, there should be a stable all purpose unified run-time configurable solution out there… but all of that sweat was squandered on building second rate rushed disposable systems… a workflow which should’ve died with 16bit computing. high level development is hardly bound by 32bit computing, and is completely unbound by 64bits.

so i think i will jump ship here, now that i’ve steered it way off course. splash

Thanks, Gdewan, that really blew me away!

:smiley:

Originally posted by Adrian:
[b] [quote]Originally posted by michagl:
i don’t think a good ‘terrain’ needs 1 million polygons to be pixel perfect on forseeable displays. i would guess that most of those triangles would be wasted power.
Display technology is moving quite fast now. You can buy a 24inch Dell LCD 1920X1200 for $1200 US. I’m sure the price will fall quickly this year like the 20inch version did last year.

It’s better that the gpu at least has the opportunity to waste some vertex processing power rather than sitting idle. Despite the wastage the brute force methods will still render more useful tris/sec that a cpu heavy approach.

The 10x faster card I was referring to was an ATI X850XT platinum capable of transforming 800M vertices/sec, compared to the 40-80M of the FX500. (I’m not familiar with the FX500 I found conflicting specs of its transform speed). I believe the X850XT sells for around 500 US.[/b][/QUOTE]1200 is quite a bit if you are not an ultra consumer. i could get 4 computers easy for that price.

guess which i would rather have.

i have the same crt monitors i’ve ever had for the past 8 years… cost about 150$ a pop then, like i figure they did for a while before, and still do pretty much… though shipping might make a significant dent compared to the flat displays… but those suckers are still damn heavy last i checked.

those tvs suck up a lot of power too… especially the plasmas… be sure to add that to the cost. add the EPA bill to it while you are at it too.

those little portable lcd screens cost a pretty penny too.

once head mounted displays are reasonable… expect lores to suddenly become all the rage.

as for 500$ card… i shelled out 250$ for the quadrofx500 (a good deal) … 400$ for the one before that… saw maybe a 2 fold speedup at best, but the features were what i was really interested in.

from your 500$ card, i would expect realisticly a 2 or 3 fold boost tops… but i don’t plan to wrangle up 500$ anytime soon. 250$ would be more my style.

with these huge batch numbers though… as far as they ring true… i guess people want bigger better geometry over quantity. thats a little sad… but maybe the bus will over take the triangles someday and throw all the numbers on their heads. the hardware manufacturers should really do their best to not force developers in any direction. it kills capacity for creative growth.

edit: this last thought must’ve come out of the blue. in response though, of course developers are going to improve the hardware however they can i guess. but something like favoring dense geometry over dispersed geometry tends to force people into thinking in one way… and might even adversely effect hardware developmen trends. of course if those numbers are to be taken seriously, it would lead to the conclusion that it is equally as costly to render 1000 triangles to a pixel as it is to render 1… with those numbers any attempt at LOD seems futile. but this is simply not reality… because a system withou LOD will come to a halt instantly. try working with large models in a 3D modeling environment like Maya. there is no real geometric LOD there, and it will crash pretty quickly. i figure display lists would could take whatever driver overhead there is to batch processing… but not if you have to regenerate the list more often than not… so complex particles, like maybe Macross style space debris could not be done unless most of the particles shared transform matrices.

if the player could save in the field, the system has to be prepared to spit those 1 million polygons in view out immediately… and then grab its ass and do its best to stream the polygons which will soon come into view in asap.
I’m not sure I follow your logic here.

If you’re reloading the game (after a previous save), a hard-load operation is expected. Since the data isn’t there already, you pop up a “Loading…” screen and load it. There is no problem, and you don’t need to force the player into specific locations to save.

the difference between continuous and non-continuous LOD, which is the only remaining question for terrain, is that non-continuous LOD relies on artificially inflating the triangle count to try to obscure seams.
No, the primary performance difference between them is that static LOD gives the card the data as it wants it (large, contiguous streams), while the continuous LOD scheme is unable to do so.

The static scheme is not streaming data up to the card very often; maybe once every 5 seconds at maximum player speed. Not only is this eventuality rare, it also happens as the card would prefer: long, continuous streams of data. This is benifitial for the CPU and the GPU.

The continuous scheme has to frequently update the data, possibly in large quantities, but possibly in small portions. In either case, it hurts more on the CPU and the GPU side.

hell most code bases are built to be disposable… this is so wasteful, because if people have half the sense they like to think they do, we shouldn’t still be compiling games in 2005, there should be a stable all purpose unified run-time configurable solution out there…
I’m not sure you understand the impossibility of what you’re considering. There is no one-size-fits-all approach. The reason codebases are “built to be disposable” is because constant hardware changes and competition requires it. You can’t predict the future, and everyone who has tried to in this industry has been burned by it. People who thought that non-programmatic hardware was the future built engines around that, and these engines needed almost total rewrites because the future didn’t turn out as they hoped.

The needs of an RTS game are not the needs of an FPS game; indeed, their needs in terms of terrain, field-of-vision, and all manor of other things are very different. They have some basic needs in common (they draw meshes if it is a 3D RTS, etc), but there are plenty of needs that they do not have in common, and those games have little in common with something like GTA or Zelda. As such, running one kind of game on the engine of another, while possible, is not terribly wise. You will never make a generalized engine that can run any kind of game in any kind of genre nearly as well as you can if you made a specialized system.

More importantly, the present exists. If I knew for a fact that programmability on GPUs was coming in 2 years, would I build my engine for a game that was coming out in 1.5 years with programmability in mind? Of course not; it makes no sense. It serves no purpose to the present; the needs of a progammability-aware engine are not the same as one that isn’t. The purpose in building such a system for a non-programmability-aware GPU makes no sense, and can even make it more difficult to get that game out in time. The future isn’t here yet, so there’s no need to go to drastic lengths to plan for it.

That’s not to say that you should write rediculously short-seighted code either. Flexibility can be built into systems without going too far. Programmer prefer to build flexible systems that can be “easily” replaced by others if they no longer are appropriate for the new task. This does not mean building a gigantic monolithic system that does everything.

Maybe someday ROAM and other similar algorithms will be the right way to go. Maybe you even do ROAM-type stuff on the GPU completely. In which case, all your work will have meaning. However, for today, it is the slower path. For today, it has only limitted applications, and few of them are high-performance. For today, there are better alternatives. And if we ignore today just because we believe that things will be different tomorrow, we lose out to those who are living in the present rather than the future.

Just because there’s a cliff 1 mile in front of you doesn’t mean you turn; you still have 5,279 feet before that cliff becomes an issue.

i guess people want bigger better geometry over quantity.
I don’t know what that means. Bigger, better geometry means more (re: quantity) of geometry. Clearly, the size of a trangle has no impact on how long it takes in terms of vertex processing, so the performance benifit of using static schemes is in having more triangles.

the hardware manufacturers should really do their best to not force developers in any direction. it kills capacity for creative growth.
With that kind of logic, we’d still be back in GeForce 1-levels of performance.

You have to optimize, and you optimize where it counts. And if it means that large strings of triangles are the optimal path, so be it; at least we have some path that is optimal. The fact that it happens to be brute force (the path that hardware tends to favor) only makes it easier to use.

yo Korval,

sorry, but you misunderstood me on most every point in your last post… and i take most of the blame in that if not all of the blame. when i get the time i will respond to your words though, if only to clarify my original intent… which probably wasn’t as explicit as i would’ve liked it to be in a perfect world.

i really apreciate the concern though.

sincerely,

michael

If you’re reloading the game (after a previous save), a hard-load operation is expected. Since the data isn’t there already, you pop up a “Loading…” screen and load it. There is no problem, and you don’t need to force the player into specific locations to save.

my point is simply that for 1 million triangles on screen, unless there is some development in disk technology i’m unaware of, some serious ‘hard’ load times are going to be in order. personally if i have to look at a loading screen for more than 10 seconds, i’m bothered… i’d just assume do something else while its loading… it breaks peoples lives up into little un retrievable chunks, kind of like commuting and commercial tv. i don’t play too many video games, if any by any measurable standard. but the ones i do have, have rediculous loading procedures, which are generally unecesarry, like reloading the environment you were just in for a replay or something. i would probably play more often if not for that. the load times these days are just rediculous, but if people are accustomed to it, i guess they get what they ask for.

No, the primary performance difference between them is that static LOD gives the card the data as it wants it (large, contiguous streams), while the continuous LOD scheme is unable to do so.

The static scheme is not streaming data up to the card very often; maybe once every 5 seconds at maximum player speed. Not only is this eventuality rare, it also happens as the card would prefer: long, continuous streams of data. This is benifitial for the CPU and the GPU.

The continuous scheme has to frequently update the data, possibly in large quantities, but possibly in small portions. In either case, it hurts more on the CPU and the GPU side.

just curious if uploading data halts the cpu in current implimentations if the app is not multi-threaded? this is something i don’t know much about. and just for the record the algorithm discussed here only uploads continuously, but large has yet to be seen. the thing that scares me about large uploads, isn’t the uploading. its where to get all of the data to upload. for a nurbs surface for instance, all of the vertices must be procedurally generated, but reading from disk is not much better right now with cpu trends versus disk io. any LOD algorithm which utilizes ‘mipmapping’ to generate displaced geometry must recompute the mesh for each level, or risk non-representative sampling by not utilizing mipmapping, which would mean more visual popping in the LOD. so i have to ask if we are talking about stream non-continuous LOD here, or just streaming in a static mesh with no LOD in chunks?

I’m not sure you understand the impossibility of what you’re considering. There is no one-size-fits-all approach. The reason codebases are “built to be disposable” is because constant hardware changes and competition requires it. You can’t predict the future, and everyone who has tried to in this industry has been burned by it. People who thought that non-programmatic hardware was the future built engines around that, and these engines needed almost total rewrites because the future didn’t turn out as they hoped.

there are reasonable ways to go about providing hardware hooks without throwing the baby out with the bath water… assuming the baby is worth keeping at all in the first place. producing low quality systems has just become habitual. its a waste of energy… but i can understand why it has flurished, because building high quality systems is hard to coordinate with large teams, especially large teams who do not actually rely on the applications they develope – though that is less of a problem in the video game industry i figure.

The needs of an RTS game are not the needs of an FPS game; indeed, their needs in terms of terrain, field-of-vision, and all manor of other things are very different. They have some basic needs in common (they draw meshes if it is a 3D RTS, etc), but there are plenty of needs that they do not have in common, and those games have little in common with something like GTA or Zelda. As such, running one kind of game on the engine of another, while possible, is not terribly wise. You will never make a generalized engine that can run any kind of game in any kind of genre nearly as well as you can if you made a specialized system.

yeah of course, different systems would be used for a typical top down RTS versus a first person type configuration… same as say a largely 2D verus a 3D configuration. the trick is just to paint as wide a swatch as possible as far as abstraction is concerned, without introducing overburdening dependancies.

More importantly, the present exists. If I knew for a fact that programmability on GPUs was coming in 2 years, would I build my engine for a game that was coming out in 1.5 years with programmability in mind? Of course not;

no you’d be wise to leave hardware to hardware and software to software. that is keep the two seperate dependancy wise. i would personally never build hardware into a system directly. you are right though that graphics hardware 2 or 4 years ago was a very chaotic affair. i think it is fairly well stabilizing though to the point that its behavior can be fairly well predicted and abstracted with the introduction of the programmable gpu… in my experience 64bit processing allows for a much more advantageous environment for recoverable development than 32bit really allowed for… though i believe it could’ve and should’ve been done with 32bit. the window is very wide for 64bit though… and the restrictions of 32bits on matters such as run-time programming (lisp), precision, and addressing will not be present with 64bits. and i don’t expect 128bit processing anytime soon… i think there is a noticible curve going from the amount of time to transission from 8 to 16 to 32 to 64 bits, which probably correlate with the order of magnitude of the values achievable by such alignments. in other words, developments are becoming increasingly stabilized in the big picture.

it makes no sense. It serves no purpose to the present; the needs of a progammability-aware engine are not the same as one that isn’t. The purpose in building such a system for a non-programmability-aware GPU makes no sense, and can even make it more difficult to get that game out in time. The future isn’t here yet, so there’s no need to go to drastic lengths to plan for it.

if you can save time on the next iteration it is worth while… especially if you invested in the last cycle… and i’m sure this is done to a degree.

That’s not to say that you should write rediculously short-seighted code either. Flexibility can be built into systems without going too far. Programmer prefer to build flexible systems that can be “easily” replaced by others if they no longer are appropriate for the new task. This does not mean building a gigantic monolithic system that does everything.

yeah i agree… now is a good time i think to start shifting in this direction. as the days pass by, this approach becomes more and more reasonable. the trick is just being able to recognize and admit that the old ways are counter productive when the times come and the environments change.

the real trick with monolithic systems though i think is, throwing more people at them does no good. so you need a sort of ‘heroic’ programmer or two who can think with the same mind to pull it off, or at least get it started.

Maybe someday ROAM and other similar algorithms will be the right way to go. Maybe you even do ROAM-type stuff on the GPU completely.

i suspect it will go completely on hardware. the question i think isn’t a matter of ‘if’, but ‘when’. would work well probably with the regular mesh images advocated by ‘hoppe’, and nurbs parameters. but in the mean time i don’t believe it is inefficient to pull off with contemporary hardware capacities… just takes a lot more thought to impliment.

In which case, all your work will have meaning. However, for today, it is the slower path. For today, it has only limitted applications, and few of them are high-performance. For today, there are better alternatives. And if we ignore today just because we believe that things will be different tomorrow, we lose out to those who are living in the present rather than the future.

i’m not advocating anything… the results will speak for themselves if i have anything to say for it. the only strong counter argument i see is this whole batching thing which proports that a single glDrawElements dispatch sent across agp memory is as fast as drawing 1000 simple triangles. first of all this breaks down when your shaders have enough instructions… notice how coarse the geometry is with modern games. but even for simple shaders, i have a feeling utilizing displaylists will totally rip the performance hit out of that batch argument. if you have doubts about the cpu overhead for my ROAM algorithm, i have a feeling they are unfounded, because the cpu overhead is nothing.

Just because there’s a cliff 1 mile in front of you doesn’t mean you turn; you still have 5,279 feet before that cliff becomes an issue.

this depends on just how fast you age going :slight_smile:

for my own needs, i’m going very fast… or if you don’t like the ‘fast’ analogy… then lets just say my reaction time is very slowwww.

I don’t know what that means. Bigger, better geometry means more (re: quantity) of geometry. Clearly, the size of a trangle has no impact on how long it takes in terms of vertex processing, so the performance benifit of using static schemes is in having more triangles.
yeah, i was just commenting, that just due to the bus versus the rasterization hardware… current trends favor an environment with a few large high resolution objects… versus even a lot of small high resolution objects. say we can’t have 1000 rocks rolling down a ravine, we have to have a single boulder. i’m not saying that is anyone’s fault necesarrilly, just unfortunate. however if people grow too accustomed to this thinking, they might pass up an opertunity to improve the bus while focusing too much on the rasterizer. tunnel vision or something.

With that kind of logic, we’d still be back in GeForce 1-levels of performance.

You have to optimize, and you optimize where it counts. And if it means that large strings of triangles are the optimal path, so be it; at least we have some path that is optimal. The fact that it happens to be brute force (the path that hardware tends to favor) only makes it easier to use.
yeah i agree, that any improvement is improvement… but if resources are just being focused along one path versus another, i think the aproach is lop sided. and only favors competitive mediocrity.