6115
Comment:
|
7911
|
Deletions are marked like this. | Additions are marked like this. |
Line 1: | Line 1: |
#rev 2018-08-27 mreimers #rev 2020-08-31 alders |
|
Line 13: | Line 16: |
<<Anchor(2021-07-05-linux-printing)>> == Linux printing affected by PrintNightmare vulnerability patch == '''Status:''' {{attachment:Status/red.gif}} 2021-07-05 09:41:: Authentification for printing fails. Ticket at ID servicedesk opened. |
|
Line 14: | Line 21: |
<<Anchor(2017-10-24-ibtnas02)>> | <<Anchor(2021-04-27-itetmaster01-update)>> == Downtime various D-ITET services for server maintenance == '''Status:''' {{attachment:Status/green.gif}} |
Line 16: | Line 25: |
== Outage of Server ibtnas02 == '''Status:''' {{attachment:Status/orange.gif}} |
2021-04-27 08:30:: Condor is back online, all services restored. 2021-04-27 08:15:: Matrix/Element Chat services back online. 2021-04-27 08:00:: Database upgrade done and online. 2021-04-27 07:30:: Slurm services are back online. 2021-04-27 07:00:: Base system has been upgraded, main database services in progress of upgrade. 2021-04-27 06:00:: On 2021-04-27 between 06:00 and 08:30 ISG is going to update a server providing access to various D-ITET services. During the migration the following services will be affected and offline: * Matrix/Element Chat services (the instances will be unavailable) * IFA/Control Website: Access to the IFA database is blocked * Slurm (D-ITET Arton Cluster): It won't be possible to submit new jobs or view Slurm statistics. Already running jobs will not be affected. * Condor: Condor clients will be shut down the evening before to avoid running jobs during the migration. |
Line 19: | Line 36: |
2017-10-25 10:00:: ibtnas03 now serves all partitions but the problem is not yet identified | <<Anchor(2021-03-31-network-disruption)>> == Network disruption affecting several ISG.EE services == '''Status:''' {{attachment:Status/green.gif}} |
Line 21: | Line 40: |
2017-10-24 15:00:: The server ibtnas03 is up again ( partition data-08 is not available) | 2021-03-31 09:30:: The configuration error was found. The configuration change will be deployed on '''2021-04-01 around 06:15''' and a short network of about 1min is expected. 2021-03-31 08:00:: ID Networking team has rolled-back a deployed configuration, pending further investigation/analysis. 2021-03-31 07:30:: There are currently disruption affecting a VPZ with servers managed by ISG.EE. Networking team of ID is investigating the issue. There are several ISG.EE services affected/malfunctioning due to this in particluar the FindYourData service. |
Line 23: | Line 44: |
2017-10-24 12:50:: The server ibtnas03 is down again | <<Anchor(2021-03-11-mira-upgrade)>> == login.ee.ethz.ch: downtime for server upgrade == '''Status:''' {{attachment:Status/green.gif}} |
Line 25: | Line 48: |
2017-10-24 09:30:: The server ibtnas03 is back online | 2021-03-11 06:30:: Upgrade completed and service is up and running again. 2021-03-11 06:00:: The server servicing login.ee.ethz.ch will be upgraded to a new OS version (Debian buster). During the time of the update logins might not be possible. |
Line 27: | Line 51: |
2017-10-24 08:00:: The server ibtnas03 is down due to hardware problems | <<Anchor(2020-07-11-storage-downtime)>> == Planned project/ archive storage downtime and client reboot == '''Status:''' {{attachment:Status/green.gif}} |
Line 29: | Line 55: |
<<Anchor(2017-10-21-itetnas03)>> | 2020-07-11 12:00:: Migration has been completed, all services are back to operational state. |
Line 31: | Line 57: |
== Outage of Server itetnas03 == '''Status:''' {{attachment:Status/green.gif}} 2017-10-23 18:15:: Data are also accessible via NFS. |
2020-07-11 08:00:: Migration started, services are shutdown |
Line 36: | Line 59: |
2017-10-23 9:30:: The server is up. Data are accessible via Samba. NFS file service is still down. 2017-10-21 15:00:: The server itetnas03 is down due to hardware problems |
2020-07-11 8:00-12:00:: Start of planned maintenance work. Project/ archive storage services (known under the names "ocean", "bluebay", "lagoon" and "benderstor") will not be available. ISG-managed Linux clients will be rebooted. |
Line 42: | Line 63: |
<<Anchor(2017-10-18-outage-etz-d-96-2)>> == Outage of Servers in Serverroom ETZ/D/96.2 == '''Status:''' {{attachment:Status/green.gif}} 2017-10-20 13:45:: All racks in ETZ/D/96.2 are working again (cooling problem solved). 2017-10-20 10:00:: The technician will arrive at 13:00 hours. Some servers are running, but without watercooling. So any rack might shutdown at any time if the air cooling is not sufficient. This will most probably again happen when the technician will be working in the room (i.e. this afternoon). 2017-10-19 18:30:: The cooling engineer could not fix the problem, so some servers are still offline. Another technicial will try to fix the cooling system tomorrow morning. 2017-10-18 14:00:: Cooling system is still not working correctly, we only selectively powered on a couple of compute machines. 2017-10-18 12:50:: The problem has been localized and repaired. We need to wait that the circuit is cooling down. 2017-10-18 10:30:: Outage of most racks in ETZ/D/96.2 (cooling problem) . Most compute servers are offline. <<Anchor(2017-05-13-outage-etz-d-96-2)>> == Outage Servers in Serverroom ETZ/D/96.2 == |
<<Anchor(2020-06-04-svnsrv-upgrade)>> == svn.ee.ethz.ch downtime for server upgrade == |
Line 65: | Line 67: |
2017-05-13 20:00:: Outage of some racks in ETZ/D/96.2. Several compute servers offline. 2017-05-13 23:59:: Most of the servers are back online. 2017-05-15 08:45:: Status of remaining servers verified. All back online. |
2020-06-04 07:05:: Webservices for managing SVN repositories are enabled. 2020-06-04 06:15:: Systemupgrade is done and access to the SVN repositories via the `svn` and `https` transport protocols are back online. 2020-06-04 06:00:: The server servicing the SVN repositories will be upgraded to a new operating system version. During this timeframe outages for access to the SVN repositories are expected. |
Line 69: | Line 71: |
<<Anchor(2017-03-24-cronbox-login-ssh-keys)>> | <<Anchor(2020-05-17-cluster-abuse)>> == European HPC cluster abuse == '''Status:''' {{attachment:Status/green.gif}}<<BR>> Recently European HPC clusters have been attacked and abused for mining purposes. The D-ITET Slurm and SGE clusters have not been compromised. We are monitoring the situation closely. 2020-05 17 08:30:: No successful login from known attacker IP addresses could be determined, none of the files indicating being compromised have been found on our file systems 2020-05-16 14:30:: No unusal cluster job activity was observed |
Line 71: | Line 78: |
== Cronbox/Login Server migration: new SSH host key == | <<Anchor(2020-05-04-itetnas04-upgrade)>> == D-ITET Netscratch downtime for server upgrade == |
Line 74: | Line 82: |
2017-03-24 17:00:: The cronbox and login server has moved to a new host. A new SSH host key has been generated: {{{ 4096 MD5:fc:a8:00:5b:64:90:86:a1:fb:49:75:ef:55:58:90:b3 (RSA) 4096 SHA256:v48HAAAjr+avnPAESdQzazSriKYZeTGGtIPKfoE8Dg0 (RSA) }}} The SSH host key is as well listed on: https://people.ee.ethz.ch/ |
2020-05-04 06:00:: Server upgrade has been completed. 2020-05-04 06:00:: The server servicing the D-ITET Netscratch service will be upgraded to a new operating system version. During this timeframe outages for the NFS service will be expected. |
Line 81: | Line 85: |
Remember:: '''Always''' verify a fingerprint of a SSH host key before accepting it. | <<Anchor(2020-04-07-network-interuption)>> == Network outage ETx router == '''Status:''' {{attachment:Status/green.gif}} 2020-04-07 05:30:: There was an issue on the Router `rou-etx`. ID networking team trackled and solved the issue. There was about a 10min interuption for the ETx networking zone affecting almost all ISG.EE maintained systems. |
Line 83: | Line 90: |
<<Anchor(2017-01-07-Mailsystem migration)>> | <<Anchor(2020-04-06-mira-maintenance)>> == login.ee.ethz.ch: Reboot for maintenance == '''Status:''' {{attachment:Status/green.gif}} 2020-04-06 05:35:: System behind `login.ee.ethz.ch` has been rebootet for maintenance and increase available resources. |
Line 85: | Line 95: |
== EE Mailsystem migration == '''STATUS:''' {{attachment:Status/green.gif}} '''Mailsystem up''' |
See the [[RemoteAccess|information on access D-ITET resources remotely]]. To distribute better the load user are encouraged to use the VPN service whenever possible. |
Line 88: | Line 97: |
2017-01-08 15:00:: The new mailsystem is now started. In case of unattended problems we will stop it again to prevent data loss and to analyze the problem. | <<Anchor(2020-02-18-nostro-maintenance)>> == itet-stor (FindYourData) Server maintenance: Reconfiguration of VM parameters == '''Status:''' {{attachment:Status/green.gif}} |
Line 90: | Line 101: |
2017-01-07 24:00:: Not all testcases could be performed. We now plan to enable the new system about noon. | 2020-02-18 19:03:: System again up and running. 2020-02-18 19:00:: Scheduled downtime for the [[Workstations/FindYourData|itet-stor/FindYourData service]] due to maintenance work on the underlying server. |
Line 92: | Line 104: |
2017-01-07 20:45:: Old Mailserver Configuration migrated, starting the mailserver testing | <<Anchor(2020-01-20-nostro-os-upgrade)>> == itet-stor (FindYourData) Server migration: New operating system version == '''Status:''' {{attachment:Status/green.gif}} |
Line 94: | Line 108: |
2017-01-07 14:00:: User mailbox data migrated, starting mailserver configuration migration | 2020-01-20 07:15:: OS upgrade done, there were short interruptions to the [[Workstations/FindYourData|itet-stor/FindYourData service]]. 2020-01-20 06:00:: We will update the server servicing the [[Workstations/FindYourData|FindYourData service]] from Debian jessie 8 to Debian stretch 9. There will be short downtimes accessing this service during the update. |
Line 96: | Line 111: |
2017-01-07 07:00:: All mail services are stopped. Mailbox data copy started. <<Anchor(2016-09-12-network-outage)>> == Networkoutage ETH == '''STATUS:''' {{attachment:Status/green.gif}} 2016-02-09 08:20:: ETH wide network outage due to hardware problems for the firewall infrastructure. In any case, please reboot your computer before continue. 2016-02-09 12:35:: Network is back online and services are being recovered. Due to the hardware failure 53 network zones were affected. The problem got localized and resolved. 2016-02-09 14:25:: Our systems should be all back to normal. In case you experience any problem please contact support via mailto:support@ee.ethz.ch. <<Anchor(2016-02-10-maintenance-polaris)>> == Maintenance login.ee.ethz.ch and cronbox.ee.ethz.ch service == '''STATUS:''' {{attachment:Status/green.gif}} 2016-02-10: 06:05:: The server for the [[Services/Cronjob|cronbox]] and login service is currently beeing updated from Debian Wheezy to Debian Jessie. The services will be temporarly unavailable. 2016-02-10: 12:00:: Server update is done. |
|
Line 120: | Line 114: |
[[Status/Archive/2010|2010]] [[Status/Archive/2011|2011]] [[Status/Archive/2012|2012]] [[Status/Archive/2013|2013]] [[Status/Archive/2014|2014]] |
|
Line 121: | Line 120: |
[[Status/Archive/2014|2014]] [[Status/Archive/2013|2013]] [[Status/Archive/2012|2012]] [[Status/Archive/2011|2011]] [[Status/Archive/2010|2010]] |
[[Status/Archive/2016|2016]] [[Status/Archive/2017|2017]] [[Status/Archive/2018|2018]] [[Status/Archive/2019|2019]] |
General Informations
This page lists announcements and status messages for IT services managed by ISG.EE.
For notifications and announcements of central IT services managed by ID, please visit https://www.ethz.ch/services/de/it-services/service-desk.html
For a detailed status overview of central IT services managed by ID, please visit https://ueberwachung.ethz.ch
Status-Key |
|
|
Resolved |
|
Still working but with some errors |
|
Pending |
Current status reports
Linux printing affected by PrintNightmare vulnerability patch
Status:
- 2021-07-05 09:41
- Authentification for printing fails. Ticket at ID servicedesk opened.
Downtime various D-ITET services for server maintenance
Status:
- 2021-04-27 08:30
- Condor is back online, all services restored.
- 2021-04-27 08:15
- Matrix/Element Chat services back online.
- 2021-04-27 08:00
- Database upgrade done and online.
- 2021-04-27 07:30
- Slurm services are back online.
- 2021-04-27 07:00
- Base system has been upgraded, main database services in progress of upgrade.
- 2021-04-27 06:00
- On 2021-04-27 between 06:00 and 08:30 ISG is going to update a server providing access to various D-ITET services. During the migration the following services will be affected and offline:
- Matrix/Element Chat services (the instances will be unavailable)
- IFA/Control Website: Access to the IFA database is blocked
- Slurm (D-ITET Arton Cluster): It won't be possible to submit new jobs or view Slurm statistics. Already running jobs will not be affected.
- Condor: Condor clients will be shut down the evening before to avoid running jobs during the migration.
Network disruption affecting several ISG.EE services
Status:
- 2021-03-31 09:30
The configuration error was found. The configuration change will be deployed on 2021-04-01 around 06:15 and a short network of about 1min is expected.
- 2021-03-31 08:00
- ID Networking team has rolled-back a deployed configuration, pending further investigation/analysis.
- 2021-03-31 07:30
There are currently disruption affecting a VPZ with servers managed by ISG.EE. Networking team of ID is investigating the issue. There are several ISG.EE services affected/malfunctioning due to this in particluar the FindYourData service.
login.ee.ethz.ch: downtime for server upgrade
Status:
- 2021-03-11 06:30
- Upgrade completed and service is up and running again.
- 2021-03-11 06:00
- The server servicing login.ee.ethz.ch will be upgraded to a new OS version (Debian buster). During the time of the update logins might not be possible.
Planned project/ archive storage downtime and client reboot
Status:
- 2020-07-11 12:00
- Migration has been completed, all services are back to operational state.
- 2020-07-11 08:00
- Migration started, services are shutdown
- 2020-07-11 8:00-12:00
- Start of planned maintenance work. Project/ archive storage services (known under the names "ocean", "bluebay", "lagoon" and "benderstor") will not be available. ISG-managed Linux clients will be rebooted.
svn.ee.ethz.ch downtime for server upgrade
Status:
- 2020-06-04 07:05
- Webservices for managing SVN repositories are enabled.
- 2020-06-04 06:15
Systemupgrade is done and access to the SVN repositories via the svn and https transport protocols are back online.
- 2020-06-04 06:00
- The server servicing the SVN repositories will be upgraded to a new operating system version. During this timeframe outages for access to the SVN repositories are expected.
European HPC cluster abuse
Status:
Recently European HPC clusters have been attacked and abused for mining purposes. The D-ITET Slurm and SGE clusters have not been compromised. We are monitoring the situation closely.
- 2020-05 17 08:30
- No successful login from known attacker IP addresses could be determined, none of the files indicating being compromised have been found on our file systems
- 2020-05-16 14:30
- No unusal cluster job activity was observed
D-ITET Netscratch downtime for server upgrade
Status:
- 2020-05-04 06:00
- Server upgrade has been completed.
- 2020-05-04 06:00
- The server servicing the D-ITET Netscratch service will be upgraded to a new operating system version. During this timeframe outages for the NFS service will be expected.
Network outage ETx router
Status:
- 2020-04-07 05:30
There was an issue on the Router rou-etx. ID networking team trackled and solved the issue. There was about a 10min interuption for the ETx networking zone affecting almost all ISG.EE maintained systems.
login.ee.ethz.ch: Reboot for maintenance
Status:
- 2020-04-06 05:35
System behind login.ee.ethz.ch has been rebootet for maintenance and increase available resources.
See the information on access D-ITET resources remotely. To distribute better the load user are encouraged to use the VPN service whenever possible.
itet-stor (FindYourData) Server maintenance: Reconfiguration of VM parameters
Status:
- 2020-02-18 19:03
- System again up and running.
- 2020-02-18 19:00
Scheduled downtime for the itet-stor/FindYourData service due to maintenance work on the underlying server.
itet-stor (FindYourData) Server migration: New operating system version
Status:
- 2020-01-20 07:15
OS upgrade done, there were short interruptions to the itet-stor/FindYourData service.
- 2020-01-20 06:00
We will update the server servicing the FindYourData service from Debian jessie 8 to Debian stretch 9. There will be short downtimes accessing this service during the update.
Archived status reports
2010 2011 2012 2013 2014 2015 2016 2017 2018 2019