Interactive sessions can't connect after idle

We’re running version 3.1.16 and having a similar problem to this older topic.

Our users launch an interactive session, step away, and then when they try to reconnect they see a VNC error rejecting the authentication:

I can sometimes replicate this simply by disconnecting from the session and reloading the browser window. Then when I try to reconnect I see this error. The “view only” link works to connect to the session and see what is running.

I’ve been unable to find a fix for this and the users have to cancel their jobs and start over. Does anyone have any suggestions?

Thanks,
Dori

Sorry for the delay in response Dori.

When you say “they try to reconnect” are you saying they try to connect on the same page/refresh that page or are they trying to reconnect through the button Launch <cluster> Desktop>? The former won’t work because these passwords are one time use. Every time you connect using a given password, it regenerates a new one for the next time. The latter should continue to work because the card should be able to pick up the newer password.

When I refresh the page once, I just get authentication failed. But if I continue to refresh the page - I get that message - authentication failed. client temporarily disabled.

However if I navigate back to OnDemand, the Launch <cluster> Desktop button still works using the new password.

So I think it may be a user error - refreshing that page won’t work, but using the Launch <cluster> Desktop button in OnDemand should continue to work.

If that button stops working, let me know and we’ll continue to look at it.

Hi Jeff,

No problem! We were looking into things on our end too. I am able to replicate this on demand (no pun intended!) now. I start up a session (desktop, app, doesn’t matter) and click the button to launch it. Once that new tab or window with the session is open, it works. If I close the tab and click the button to launch it again, I get the connection error. We thought it might be a firewall issue but confirmed in our dev environment that isn’t the case. There isn’t anything in the user nginx logs or the job logs indicating a problem. Are there other logs we could be looking at that might indicate what might be causing the problem?

Thanks,
Dori

Interesting.

You should see messages like this in your output.log as you generate new passwords. Every time you successfully connect it should regenerate a new password. Maybe you’re clicking the card too soon as it hasn’t picked up the new password yet? Or maybe there’s some NFS lag where the file is being updated on the compute node, but the web node is still reading the old file?

When I click the button on our 4.2 instance, it disables it for a bit so I can’t click it that fast.

(polkit-mate-authentication-agent-1:3387467): polkit-mate-1-WARNING **: 10:57:01.157: Unable to determine the session we are in: No session for pid 3387467
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...

Or maybe the page is failing to update the card. If you open your network tab on your ‘my interactive sessions’ page you should see something like this where the browser is querying the backend for new information (including a new password).

Hi Jeff,

Here’s an update on what we’ve seen and tested.

  • Doesn’t appear to be browser based. It’s happening in Chrome, Firefox, and Edge
  • I can’t replicate it immediately with interactive apps, including VNC based ones like the Matlab app. It’s only desktops. Eventually I can get the Matlab app to break by closing and reopening the browser tab with the session a few times.
  • When it’s working, I’m seeing the output.log file indicate that a new VNC password is generated every time I click the button to launch the app. When it’s not working, I only see this with the initial launch of the desktop but not afterwards
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
Setting VNC password...
Generating connection YAML file...
  • The non-VNC apps (Rstudio, Jupyter) work fine with no issues. Even if the session times out, the user can click the button and get a new one.
  • The view-only links work fine for all the apps

We’re not seeing anything in the user logs, job logs, or the network info in the browser. We do have plans to update to 4.x in the next few months but aren’t prepared to do that right now. This is causing an issue for users that start something running in their session and leave it. Then they can’t connect back to it and have to resubmit their jobs which is a big issue for our GPU users who wait a long time to get a node. We’re just unsure what else we can check or what could be causing this.

Thanks for any help you can provide!

Dori

This seems to be the biggest clue. This is how this is supposed to work

I wonder if you can share the full output.log and vnc.log of a desktop that stopped working?

I can’t quite imagine why it would stop working, but maybe that back-grounded tail process stopped for some reason?

I’m trying to replicate now, but you should be able to find a tail process like this in the users’ job in question.

[johrstrom ~()] 🐸  ps -elf | grep johrstr | grep tail
0 S johrstr+ 3663189 3663184  0  80   0 - 55252 hrtime 09:15 ?        00:00:00 tail -f --pid=3662923 vnc.log

Hi Jeff,

Thanks for confirming. That’s how we understood it too. I’m attaching both those files. I’m not seeing that process at all for the user.

Dori

output-log.txt (4.5 KB)

vnc-log.txt (5.1 KB)

Is it specific users or are you able to consistently replicate it?

Here’s the PID I’m waiting for in my example (3662923) - you see it’s the script.sh. Odd that the tail on the PID would fail, but I’m guessing the actual PID/script.sh is still alive?

johrstrom ~()] 🐨  ps -elf | grep johrstr | grep 3662923
0 S johrstr+ 3662923 3662849  0  80   0 - 55852 do_wai 09:14 ?        00:00:00 bash /users/PZS0714/johrstrom/ondemand/data/sys/dashboard/batch_connect/sys/bc_desktop/cardinal/output/4bec98e0-63a7-46f5-ad0c-1db5e6842f66/script.sh
0 S johrstr+ 3662996 3662923  0  80   0 - 164497 do_pol 09:15 ?       00:00:00 mate-session
0 S johrstr+ 3663189 3663184  0  80   0 - 55252 hrtime 09:15 ?        00:00:00 tail -f --pid=3662923 vnc.log

Sorry, long username was getting cut off. I am seeing that process.

ps -elf |grep dsajdak|grep tail
0 S dsajdak+  826775  826773  0  80   0 -   634 -      11:08 ?        00:00:00 tail -f --pid=826558 vnc.log

ps -elf |grep dsajdak|grep script
4 S dsajdak+  825968  825959  0  80   0 -  1208 -      11:07 ?        00:00:00 /bin/bash /var/spool/slurmd/job25946772/slurm_script
0 S dsajdak+  826558  825968  0  80   0 -  1196 -      11:08 ?        00:00:00 bash /user/dsajdaktest/ondemand/data/sys/dashboard/batch_connect/sys/bc_desktop/20-desktops/output/db22a04e-f3c0-4374-9b70-f56b54e18ae7/script.sh
1 S dsajdak+  826773  825968  0  80   0 -  1241 -      11:08 ?        00:00:00 /bin/bash /var/spool/slurmd/job25946772/slurm_script

This is happening for all users, unfortunately.

OK, we run turbovnc 3.1.1, I’m trying to build 3.1.2 to replicate. We may have to meet so I can see this interactively.

I tried to take a screen recording and I can’t replicate it consistently now. Very odd! I can even login to OOD from another browser and connect to a running session without issue. Right up until it stops working. They used to be immediately but now I have to connect a bunch of times before it finally fails. We’ll keep looking to see what might be causing it!

Do you have availability to meet later this week? I think I’d like to see it in person if you can replicate consistently, which I guess you can’t anymore. I’ll still think on it from my side.

Just tested quickly with turbovnc 3.1.2 and it seemed OK.

Thinking about this a little more - I wonder if a simple page refresh doesn’t fix it? That is when it starts happening, just refresh the ‘My Interactive Sessions’ page to pickup the new password would work.

Unfortunately, no, that doesn’t work. In fact, when I do that I can almost always get it to break.

Odd. What about the network tab in the browser when it fails. Do you see the same 200 OK’s that I’ve given above or do those requests fail?

I have tried to replicate in a 3.1 instance in development mode, but haven’t got much luck.

I’d say if you can replicate pretty consistently then we should set up a time to meet and debug it interactively since you’ve got a support subscription.