Skip to content

Wait for daemon incr/decr asynchronously - #166

Merged
edan-bainglass merged 1 commit into
aiidateam:masterfrom
edan-bainglass:async-daemon-wait
Jul 30, 2026
Merged

edan-bainglass merged 1 commit into
aiidateam:masterfrom
edan-bainglass:async-daemon-wait

Conversation

@edan-bainglass

@edan-bainglass edan-bainglass commented Jul 30, 2026 •

Copy link
Copy Markdown
Member

Closes #162

@edan-bainglass edan-bainglass changed the title Async daemon wait Wait for daemon incr/decr asynchronously Jul 30, 2026
@edan-bainglass

Copy link
Copy Markdown
Member Author

@khsrali can I get your eyes on this? 🙏

@edan-bainglass edan-bainglass self-assigned this Jul 30, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 89.06%. Comparing base (cbc1e34) to head (cb9af41).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #166   +/-   ##
=======================================
  Coverage   89.05%   89.06%           
=======================================
  Files          49       49           
  Lines        2029     2030    +1     
=======================================
+ Hits         1807     1808    +1     
  Misses        222      222           
Flag Coverage Δ
pytests 89.06% <100.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@@ -156,7 +157,7 @@ async def increase_daemon_worker() -> DaemonStatus:

initial = client.get_numprocesses()['numprocesses']

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
initial = client.get_numprocesses()['numprocesses']

initial = client.get_numprocesses()['numprocesses']
client.increase_workers(1)
num_workers = _wait_for_num_workers(initial + 1)
num_workers = await _wait_for_num_workers(initial + 1)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
num_workers = await _wait_for_num_workers(initial + 1)
num_workers = client.get_numprocesses()['numprocesses']

@Bud-Macaulay Bud-Macaulay left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi Edan, could you try avoiding using this wait_for_number... altogether.

Looking at the daemon client I can't see a reason why this wouldn't work?

Then we can drop this wait_for... etc.

@edan-bainglass

edan-bainglass commented Jul 30, 2026 •

Copy link
Copy Markdown
Member Author

Hi Edan, could you try avoiding using this wait_for_number... altogether.

Looking at the daemon client I can't see a reason why this wouldn't work?

Then we can drop this wait_for... etc.

Reducing the logic of increase_daemon_worker (which also needs to be pluralized, btw) to

async def increase_daemon_worker() -> DaemonStatus:
    """Increase the number of daemon workers by one."""
    client = get_daemon_client()

    if not client.is_daemon_running:
        raise DaemonException('The daemon is not running.')

    client.increase_workers(1)
    num_workers = client.get_numprocesses()['numprocesses']

    return DaemonStatus(running=True, num_workers=num_workers)

and pinging it consistently gives the correct (future) number of workers. However... doing this for decrease interestingly consistently gives the wrong (past) number of workers 🤔

Note that with the implementation done in this PR, I get correct counts for both ops.

@Bud-Macaulay

Copy link
Copy Markdown

Hmm interesting, looks like a bug or a race condition in teh daemon-manager/circus somewhere. I guess we can keep the repoll timeout for now but would be great to revisit this when we have a new daemon.

Merge with whatever changes you see fit.

@khsrali

khsrali commented Jul 30, 2026

Copy link
Copy Markdown

and pinging it consistently gives the correct (future) number of workers. However... doing this for decrease interestingly consistently gives the wrong (past) number of workers 🤔

Yeah, that's racing is expected, welcome to async battles😆. Perhaps because you already had send in increase request before and you are "awaiting it" elsewhere and at the same time are sending asynchronous request to decrease it. All messages gets acknowledged by circus but the number that each function call is expecting is not gonna what they've expected because it's changed by the other one.

Note this is entirely unrelated to circus being sync or async. Even if we were using async interface of circus this 100% would happen again.

The solution is to use asyncio.lock, or even better limit the number of your RESTapi calls on daemon with semaphore to 1 and exactly 1.

So as you see in the end, even an asynchronous DaemonClient if we would have implement is not gonna be as glamorous as it sounds.

@khsrali

khsrali commented Jul 30, 2026

Copy link
Copy Markdown

I guess we can keep the repoll timeout for now

Instead of repolling I suggest the solution of aiidateam/aiida-core#7500 (comment)

@edan-bainglass
edan-bainglass force-pushed the async-daemon-wait branch 2 times, most recently from cf7d1fd to 8f66382 Compare July 30, 2026 10:37
@edan-bainglass

Copy link
Copy Markdown
Member Author

Yeah, that's racing is expected, welcome to async battles😆. Perhaps because you already had send in increase request before and you are "awaiting it" elsewhere and at the same time are sending asynchronous request to decrease it. All messages gets acknowledged by circus but the number that each function call is expecting is not gonna what they've expected because it's changed by the other one.

In my testing, I was getting the wrong count from a decrease op starting fresh, so no previous calls.

Note this is entirely unrelated to circus being sync or async. Even if we were using async interface of circus this 100% would happen again.

Can you explain why in more detail?

The solution is to use asyncio.lock, or even better limit the number of your RESTapi calls on daemon with semaphore to 1 and exactly 1.

I've heard you say "semaphore" a few times now. I'm afraid I don't know what this is 🥲

@edan-bainglass

Copy link
Copy Markdown
Member Author

Instead of repolling I suggest the solution of aiidateam/aiida-core#7500 (comment)

This PR implements your suggestion from that issue. Unless you mean the recent comment you made to elevate it to aiida-core via waiting (or wait as I suggested in the issue)?

@khsrali

khsrali commented Jul 30, 2026

Copy link
Copy Markdown

This PR implements your suggestion from that issue. Unless you mean the recent comment you made to elevate it to aiida-core via waiting (or wait as I suggested in the issue)?

yes, the links goes to the latest suggestion, wait=True, which could be implemented

@edan-bainglass

Copy link
Copy Markdown
Member Author

yes, the links goes to the latest suggestion, wait=True, which could be implemented

I would be in favor of this. Just to move things along, I'm going to merge this one as is. Ping me on #159 once you guys have it implemented in aiida-core 🙏

@edan-bainglass
edan-bainglass merged commit d17ed6e into aiidateam:master Jul 30, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Asynchronously await on daemon incr/decr ops

4 participants