Efficiancy and Miner Platform proposal

Hi everyone,

For a while I’ve been trying to work out where I could be most useful to this ecosystem. It took me two overlapping drafts to see they were one program.

When asked what a Cuckatoo32 graph costs in energy and the public record offers numbers that look comparable and are not: 100 joules per graph from the iPollo nameplate, 184 from device telemetry on an A100 with the host excluded, 217 from software reported watts on an RTX 5090. The best documented Apple result has no energy figure at all. None of these share a boundary.

At today’s margins, joules per graph decides who can afford to participate, and nobody has ever put one meter behind all of it and published the answer.

That is what this program does, in two phases:
Phase I runs one frozen workload behind one external wall meter, across the Apple desktop tiers, the professional GPU candidates and a directly measured iPollo G1 Mini. Phase I closes with a straight answer to the question everything downstream depends on: can open, general purpose hardware compete, and if not yet, the exact threshold future silicon has to clear.

Phase II builds the open miner across macOS, Windows and Linux, with the Apple release guaranteed and the second backend chosen by the evidence; it activates only through a written continuation decision after the Phase I review

The infrastructure survives any verdict: the verifier, the corpora, the benchmark harness and the profiles stay public, so the next performance claim, from anyone, is cheap to check and hard to overstate.

I have attached a link to my proposal.

I would be happy to answer anyones questions.

Looking forward to hearing back from everyone!

2 Likes

Thanks for submitting this proposal and congratulations on winning the previous bounty.

Unfortunately, I haven’t yet understood the value of the proposal. In terms of knowing the power cost of mining, isn’t it a case of dividing the power use by the graph rate? So as long as you know these you can determine the power cost. In the case of an G1 mini, that would be roughly:

100 joules per graph.

120 W ÷ 1.2 graphs/s = 100 watt-seconds each.

Please let me know if I’ve misunderstood.

Thank you, and the arithmetic is right. E = P / r, and 120 W over 1.2 GPS is 100 joules per graph. The proposal uses that exact figure as its planning anchor, and it uses your formula as more than that. Section 6.1 defines the headline ranking metric as exactly that quotient, mean wall watts over resolved graphs per second; Figure 7 is the same identity rearranged, the rate each power envelope in the field would have to sustain to reach the 100 joules per graph line. The division is not in dispute; it is the program’s own instrument. What it needs is two measured inputs, and for the hardware this program is about, neither exists yet.

Apply the method to the machine the last bounty produced. The M1 Ultra miner sustains 0.8 graphs per second. To match the G1 Mini at 100 J/graph, that Mac Studio has to draw 80 W at the wall while mining. To beat the favourable end of iPollo’s own tolerance band, 81.8 J/graph, it has to hold about 65 W. Nobody knows whether it does.

Apple publishes an idle figure and a stress maximum for the chassis, and a memory-bound miner sits at an unmeasured point between them. Feed the division the 215 W stress ceiling and the Mac lands at 269 joules per graph, worse than every anchor in the field; feed it the 80 W it would need and it matches the ASIC nameplate. Same formula, same machine, opposite verdicts, the inputs are doing all the work, and the inputs are unmeasured.

There is no published wall trace for a Mac running this workload anywhere in the record. Figure 3 of the proposal lines up the public efficiency anchors, and the Apple row is the one marked “no published wall energy record.” Closing that row is the program’s first measurement.

The other anchors are taken at different boundaries, so division cannot compare them. Your 100 J/graph is whole-device at the wall. The RTX 5090’s 217 divides software-reported board power with the host computer excluded entirely. Supply conversion alone puts a floor under that exclusion: at 434 W of device draw, an 80 Plus Gold unit dissipates 38 to 48 W just delivering the power - 92 or 90 percent efficient at half load, on a 230 or 115 volt line - another 19 to 24 J/graph before the CPU, motherboard, RAM, storage or fans draw a watt.

My own A100 figure of 184 carries the same hole. A Mac Studio has no host to exclude; it is the host. The prevailing convention flatters discrete GPUs against exactly the class of hardware under evaluation. This is not a house rule of mine, either: the SPEC power methodology the protocol binds itself to cautions directly against comparing AC and DC referenced results. Division propagates a boundary mismatch. It cannot repair one.

One boundary choice the proposal makes explicit, in Section 6.1: gross whole-system energy is the headline, and incremental energy above a matched idle baseline is reported beside it as a labelled diagnostic. That is the number that matters to an owner whose machine is on anyway, and it never substitutes for the ranking figure.

The denominator is only fixed on the ASIC, and even there it moves: the rating is plus or minus ten percent on both inputs, a stock unit is reported on this forum at 0.9 to 1.1 GPS, and the community firmware held a unit at 1.31 to 1.34 for 24 hours. On general-purpose hardware, r is a property of the solver. The solver behind the 0.8 GPS figure was built to what the bounty paid for: speed, with minimal cycle loss. Energy never entered the objective function. Every engineering decision pushed toward peak throughput, and the operating points below the peak have never been swept. On a memory-bandwidth-bound workload that is where the efficient point tends to sit, because compute running flat out while it waits on DRAM is throughput-neutral and watt-expensive.

Hypothesis H1 states that principle for power-constrained GPUs, research question 6 asks it of Apple silicon, and the Month 2 plan sweeps stock, maximum-stable and best-efficiency operating points on every accessible platform. Whatever the meter reads on day one is the untuned end of that curve. And r has to be a correct r: a solver that silently drops cycles reports a flattering graph rate and a flattering J/graph, which is why the correctness and recall gates sit ahead of every energy claim.

Division answers what a G1 Mini costs to run. The funding question is whether open software on hardware people already own can compete, and there the same formula returns unknown on every input, with the general-purpose side sitting at the untuned end of its curve. A plug meter on one machine is an afternoon. A ranking the community can act on takes one boundary, a frozen workload, a physically measured iPollo in place of a nameplate, and raw traces published whichever way they fall. That is the program. If the numbers come back badly, the deliverable stands: the exact threshold future silicon has to clear, learned at the gate instead of after a build.

To be honest, I think a back-of-the-envelope calculation for mining energy use is perfectly fine.
E.g. if you run hardware, you might use the host for other things as well, the room temperatures matters, efficiency of the power supply, the maximum temperature threshold etc. I do not think we need very high precision here or that it adds much value. Rough estimates for a benchmark are fine. As such I see not much potential for a bounty personally.

As a side note, I did not open or read the PDF, they are to virus prone. Most people would prefer markdown documents or a simple forum post.

1 Like

Fair, and worth a straight answer. First, an apology on the format. The proposal runs 48 pages, with nine figures, a twenty entry claim ledger and sixty four references. It is a real proposal, and it fought at research weight, which is how it ended up typeset instead of threaded. Somewhere between the chart grids and adjusting the kerning on the cover I forgot the first rule of publishing: a document nobody opens has a readership of zero, however handsome the typography. My bad. The full text follows below in plain markdown, figures rendered as tables, and markdown first from here on.

On the substance, I agree with more of this than you might expect. High precision is not the product, and the proposal does not sell decimal places. The meter requirement exists so the instrument drops out of the argument entirely: the claims it gates are coarse ones, which machine wins, by roughly how much, and what bar the next chip must clear. Metrology is nine percent of the equipment ceiling. Precision is nearly free here, and it is not the point.
The point is that the current record is not a rough estimate. A rough estimate would be progress. For the machine the last bounty produced there is no energy number of any kind, and running the back of the envelope on what Apple does publish returns an interval that contains both 100 joules per graph, the ASIC’s nameplate, and 269 at the chassis stress maximum, which trails everything in the field. I did run that envelope. It came back reading best machine or worst machine, please measure. That is why Section 3’s anchor table marks the Apple row as the program’s first measurement.

Your list of confounders is, sincerely, the best argument for the protocol I could ask for. A host doing other work, supply efficiency, room temperature, thermal limits: effects of exactly that size are why the three public anchors cannot be ranked today. Supply conversion alone is worth about twenty joules per graph on the 5090 figure before the host draws a watt; the entire gap between the two GPU anchors is 33. Confounders that large are not something a single unmeasured estimate tolerates; they are what it hides. So the protocol treats your list as disclosure rather than noise: whole system at the wall for the ranking, incremental energy above matched idle reported beside it for the owner whose machine is on anyway, soak and drift rules for the thermals. Rough is fine. Unrankable is the thing being fixed.

And on the bounty we agree completely, which is why this is not one. A bounty pays for a single number. That instrument already ran, bought function brilliantly, and structurally cannot buy efficiency, because a hosted machine can never meet a wall meter. What is proposed is a program whose deliverables are the miner, the validated profiles, the verifier and the exact threshold future silicon has to clear, with the energy evidence acting as the gate that decides how far the build goes. If the numbers come back dull, the Council stops paying at the gate and keeps the software. Every outcome here is livable except the one where we keep debating numbers nobody has measured.

@anexus Please use your own words and not this AI garbage. Lets not “enshittify” this forum.

What, exactly do you object to? Explain the enshitification.

I mean (en)shitification in the general sence of the word and specifically how AI generated texts polutes the information space. Your reply is exactly that, very low in information density, lots of fluff.
Nor does it address the key points of critisism provided. Unecesarry fancy wording and fluff creates a future where:

Writer:
AI, Write fancy text about this simple input, make me sound smart.
Reader:
AI, decode and summarize this AI mumblejumble, its too long and fluffy to read myself.

1 Like

I think what anynomous meaning is that ,as a human being is ok to make mistake for using the words,all these mistake of words make everyone in the formun like a real person,not a machine.

I very much prefer human concise and meaningfully written text above AI beautifully sounding fluffy garbage based on a few bullet points. Honestly, AI text makes text sound nice but it is soul-les and takes extra energy to read in order to reduce all the AI filler to its original intended meaning in the form of a few bullet points.
AI has no understanding so it cannot add meaning to a text, instead it just adds beautifully sounding but meaningless filler text. AI can only work with the meaning it was provided by the user, which is often minimal and only a few bullet-points. This added filler makes text harder to read since there is no added meaning only extra work for our brain to distill whatever the original writer intended or worse, we have to ask AI to help or reduce a long text again to its original bullet points.

I would advice to read Cory Doctorow’s book “The Reverse Centaur’s Guide to Life After AI”, it very illuminating and easy to read while explaining the misconceptions and limitations and a whole lot more about AI and how we should perceive it:

2 Likes

Ya, that may have to do with poorly constructed prompts rather than a deficiency with Ai itself. Especially nowadays, agents can be trained specific examples you want to emulate so really I think the problem may be user-side.