Become a MacRumors Supporter for $50/year with no ads, ability to filter front page stories, and private forums.
At the very least many Intel Macs can now participate in the AI revolution.

💯

I think it's a bit of a shame that the 7.1 has gotten cut off at Tahoe, even if that particular macOS can be skipped. If Golden Gate has been "cleaned up" a bit with renewed efforts, I'm hoping to see a patch down the road. I much prefer to be on the latest OS for many reasons.
 
I've found the settings you mentioned above, and it totally fixed the 'maxed out' problem. Since I have decent amounts of memory (both on GPUs and system RAM), how high can I go with 'response tokens'?

And I'm guessing that the field of 'chat parameters system prompt' is where I could paste 'best practices' for coding such as the agent instructions from Paul Hudson (Hacking with Swift)?
 

Attachments

  • Skärmavbild 2026-07-07 kl. 08.39.25.png
    Skärmavbild 2026-07-07 kl. 08.39.25.png
    420.1 KB · Views: 27
Seems you missed that Golden Gate is ARM64 only, no x86. Nothing to patch.
Haha... 🤦‍♂️ *gulp* 👍🏻

I stand embarrassed, and corrected. But that makes sense. I just haven't paid it any attention at all. But then I'll just stick to Sonoma.
 
It's not possible at the moment, but I understand what you're trying to achieve. This doesn't happen with the regular text server, but it does with generation. In a future version, we'll add the functionality to use different GPUs with different ports or seeds so you can achieve this, something similar to what's available on the regular text server.

P.S I successfully added multi-instance support, which will be available in the next version, possibly later today or tomorrow.

View attachment 2643601
Hello,

I'm locked out of the 16GB models even though both of my GPUs have 16GB VRAM. Is that because of the VRAM reserve?
 

Attachments

  • Screenshot 2026-07-07 at 10.39.37 AM.png
    Screenshot 2026-07-07 at 10.39.37 AM.png
    51.4 KB · Views: 18
Hello,

I'm locked out of the 16GB models even though both of my GPUs have 16GB VRAM. Is that because of the VRAM reserve?

That might be it, but I should mention that the algorithm I implemented detects the video card's capacity along with the available space, since it also needs space for image generation. I'm going to review the algorithm today to see if there are any calculation errors; if there are, I'll let you know and release a fix in the next version.
 
  • Like
Reactions: Dolphins1972
Hello,

I'm locked out of the 16GB models even though both of my GPUs have 16GB VRAM. Is that because of the VRAM reserve?
Hi there... The fix is already applied in the latest update. I adjusted the algorithm a bit. Let me know if you can run the model correctly now. I'll be waiting for your feedback.
 
  • Like
Reactions: Dolphins1972
Hi there... The fix is already applied in the latest update. I adjusted the algorithm a bit. Let me know if you can run the model correctly now. I'll be waiting for your feedback.
Works now. Although Flux 2 generates a lot distorted images. It worked fine with your original prompt but not so good with other prompts.
Screenshot 2026-07-07 at 11.51.01 PM.png
ZIT:
Z-Image Turbo.jpg
Flux 2 Klein:
Flux 2 Klein.jpg

I forgot to mention, I'm using secondary GPU to offload the work while my main GPU remains free.
Screenshot 2026-07-08 at 12.17.26 AM.png
 
Last edited:
Tried out the instances—cool! Nice to see all four cards maxed out 🤗

A few 🪲: I see what you did with time-stamping the output images. Unfortunately for me, it seems when I click generate that whole batch of images gets time-stamped, and I will only get one output image with date and time. I see four different images in the app, but when clicking 'reveal in Finder' they all point to the same image. I guess each instance will need a unique identifier to keep them separate. But I do like the timestamps. This means that all images are kept, and not just overwritten all the time, as it was with the toshllm_output.jpg image.

Overwriting CAN be a feature, of course, to not add 100s of images over time. Perhaps a "delete/clean output images on app close" as an option?

Related to this: once I have more than one instance, after generation is completed, I don't have the 'save image' option—only reveal in Finder (with the problem above). Individual 'save image' buttons should probably be reinstated, and perhaps a 'save image set' to save all instances to a specific location as a convenience feature.

To round off for now: seeing that we'll realistically only have 2-4 instances running (?), we could order the instance grid vertically and let each instance fill more of the app's available space—more like a flex container in fill mode (as opposed to the current "thumbnail mode"). Additional information about render times and the 'save image' buttons could be to the right of each instance image. As it is now, with four instances, they all show as small thumbnails, and I think some of the info I get when only rendering one image is lost (render time and so on).

Many, many thanks for the responsive and fast development!

EDIT: depending on if it's quick fix or not: custom resolution would be nice. Great for really wide cinematic images, but I also noted that when I filled the img2img well with a reference, if that image is of a different ratio, it produces unwanted effects such as cut off top of head in portraits. So that's a tip: if you specify a 3:4 ratio for a portrait and you want to use a reference image: make sure that image is of the same proportion for best results.
 
Last edited:
Thanks for your feedback, I'll keep it in mind. I hope to resolve those UI issues better, and I'll think about the logic you requested for generating images of different resolutions because resolution is somewhat risky. If you set it to a higher resolution than your graphics card can handle, your Mac will freeze. I'll have to figure out how to handle it... perhaps as a custom aspect ratio?

P.S: You have to teach me how to create photo prompts, I'm a complete beginner at that haha
 
Thanks for your feedback, I'll keep it in mind.
Sure! I'm afraid it will just keep coming.... so with this in mind I want to underline: I'm a strong OPPONENT of "spending other people's time". It's a typical behaviour of some people: they come up with ideas and solutions, but it's always SOMEONE ELSE who has to do the work.

That's not the intention here. We already appreciate the work a lot, but as users of your app, what you'll get in this thread (that I suppose IS supposed to bring some feedback to you) is everything that either catches me off guard, or "wow, that would be nice to have...", I'll post here and then 'you do you' and decide how to best spend your time and develop the app.

Now that I've gotten that out of the way: one of the "wow, that would be nice" is prompt queuing. ESPECIALLY if the instances could be asynchronous. As it is now, once I fire up 2-4 instances and press Generate, I guess I have to wait for the whole batch to complete.

With prompt queuing, I imagine writing prompts into a queue and the next free instance just starts rendering if there are prompts queued up. That would need an output system with unique images so that they don't get overwritten, of course.

Even with a single instance, prompt queuing would be very powerful, as I've found that to get truly different results, we need to change the text prompt—the seed is not sufficient.
 
  • Like
Reactions: Flint Ironstag
P.S: You have to teach me how to create photo prompts, I'm a complete beginner at that haha

First: thanks for including the newer Flux.2 models. I realised this too late, but I would suggest changing the option of the flux.2-dev to be changed to flux.2-dev turbo (or, something newer/better, if you can find it):


The 20-50 step models do take quite a bit more time and many feel the Turbo models even surpass the larger models in many occasions. I tried the Flux.2-dev and I had to reduce output image size quite a bit in 16:9. I typically render 1600x896 but I had to go down to 1024x576.

When it comes to prompting, I've found (Z-Image Turbo—I've tried this the most) to be super straightforward. Just write plain text and ask for what you want. There is probably a skill in being explicit enough when asking for something, without creating a redundant word sallad, but I think will develop over time.

If we do manage to queue up prompts, it would be a lot easier to do research in terms of what words are needed, or whole sentences, to go a certain direction. Not sure how far you want to take this project, but in a few weeks, after summer, it would be a nice little fall project to put the Mac Pro to good use, experimenting more with local AI.
 
I'll have to figure out how to handle it... perhaps as a custom aspect ratio?
Yes, this. I shouldn't have mentioned resolution, necessarily. Custom aspect would be great. And if you're sticking to pre-defined, a proper cinematic ratio would be cool—anything from 2:1 to 2.4:1... that neighbourhood.
 
Yes, this. I shouldn't have mentioned resolution, necessarily. Custom aspect would be great. And if you're sticking to pre-defined, a proper cinematic ratio would be cool—anything from 2:1 to 2.4:1... that neighbourhood.

Haha, very late, but I'll note that last suggestion for future versions. Also, please let me know which aspect ratios to keep as the default, in addition to the custom one that's already included in the new version. I've added the prompt queues; if there's anything else to add that isn't too extensive, please let me know as well. I also need to dedicate time to trying to adapt the kernel for Vega/GCN video cards.

I'll also look into what you mentioned about the turbo... I did see that the dev version is quite large and takes much longer, so I'll see if I can integrate it later.

SCR-20260708-kmbg.png
 
Last edited:
I've added the prompt queues; if there's anything else to add that isn't too extensive, please let me know as well. I also need to dedicate time to trying to adapt the kernel for Vega/GCN video cards.

Too good! I'll take a look a bit later tonight. Life. 💪🏻🙌🏻
 
The new version has the changes you mentioned, including some cinematic aspect ratio formats. The turbo model requires a special flag that the engine doesn't yet support. For now, I've added the Klein versions, which are like the turbo version of Flux 2. I'll dedicate some time to solving the problems with the Vega cards, while I'll keep track of any issues that arise with these versions for the RDNA+ cards.
 
Crazy development, truly!

Did my first tests and it's especially nice that all the changes you implement "just work" as it were.

From a usability standpoint: I would still like to see the images in the queue occupy at least 50% of the width of that app area to see more of the result directly in the app. They are still thumbnail-esque and, at least on my monitor, there is room to spare.

It's cool to have the queue up and running, but it becomes unclear (with multiple instances running) on what instance the next queued up prompt will land. This is unfortunate, since the instances are 'pre-programmed' (perhaps even with img2img) with resolutions, ratios and such. Even different models....

What would happen if we got a 'Generate' button below each instance? Not all jobs are "single prompt with two or three seed variations" as a batch render.

As I'm thinking as I write, with a single instance the queue could work as it does now, but with multiple instances setup, we might actually need to have queue streams per instance. They could be placed next to each other in the queue interface with their respective prompts on the top and then each stream running vertically below, like now. This might take more work though?

Oh, and the queue prompt was a single line, which becomes hard to read. A textfield similar to the normal prompt, perhaps a wee bit less generous would help with prompt creation and reading.

What are your thoughts on this?

PS. I always forget to ask: do we need to update the models periodically, or does your app do that, in case of updated model releases?
 
Okay. I'll make a note of that to apply it in a future version, but I'd like a screenshot so I can better visualize how it looks on your Mac, and that way I can plan and understand it better... I hadn't thought about the models being updated; I'll see how I handle that, since adding this will require me to add the logic to detect where the model was downloaded from and check for updates, etc., from the same provider...
 
  • Like
Reactions: AndreeOnline
tried with 2xVII with no luck (frizz or dumb output)
anyway good luck

Same with my external VIIs. Haven't had a chance to try the latest model yet, but will this weekend.

Can you guys test the latest build, that includes some patchets for Vega/GCN cards, please test with and withow both Flash Attention enable and disable it... and start first with small models, and send me the logs... It should work, "should" because I still have no way of testing it directly myself; for the moment I'm going in blind and you're guiding me. It's just that, because of the earthquake, the person who was going to send me the card can't send it yet.
 
Alright...

Tried the new build and I could run Qwen image for the first time—nice. It's slow, but seems to produce quality images. Have not used it extensively. But the queue comes in handy for slower models: drop a few prompts and leave the computer to it!

Here are some screens that summarise what I tried to communicate:

toshllm_interface_1.jpg


toshllm_interface_2.jpg


toshllm_interface_3.jpg


The last images shows a mockup of my 'stretch goal'. I don't think this type of "instance queue" has any negative consequences (apart from the work to implement it of course). There could of course be a toggle close to the Generate button: 'split queue per instance'. That would make the "feature" optional for people who prefer a single queue.

Something I didn't mock up, because it's much more than a small fix, but it crossed my mind since I have 4 GPUs: I would happily dedicate one GPU to the chat and have that running, while still being able to use three instances of image generation. As it is now, we have to stop the chat, so it's either or. Not sure if that is a technical limitation, or a resource problem.

If anything is not clear, or if you want help with specific things, just let us know. I think we owe you that... 🤗
 
  • Like
Reactions: engeldlgado
Excellent, that gives me a good idea of what you want, and it doesn't seem bad at all. If there's no problem, I'll implement it next week, since I'm currently focused on the Vega cards... Thanks for the excellent suggestion!
 
  • Like
Reactions: AndreeOnline
Excellent, that gives me a good idea of what you want, and it doesn't seem bad at all. If there's no problem, I'll implement it next week, since I'm currently focused on the Vega cards... Thanks for the excellent suggestion!
I see you've been quite busy and I appreciate all the work you're doing. When you get the Vega cards figured out, can you look into enabling multi-GPU for images? I'm limited to 16GB models using only one GPU.
 
Register on MacRumors! This sidebar will go away, and you'll see fewer ads.