sonic-buildimage

Author	SHA1	Message	Date
Lawrence Lee	84cd0e9471	[mux]: Initialize all mux ports as standby Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2021-11-10 18:54:33 -08:00
Tamer Ahmed	b8f70f8986	Merged PR 3845699: [linkmgrd]: Introduce MUX cable linkmgrd Linkmgrd monitors link status, mux status, and link state. Has the link becomes unhealthy, linkmgrd will trigger mux switchover on a standby ToR ensuring uninterrupted service to servers/blades. This PR is initial implementation of linkmgrd. Also, docker-mux container hold packages related to maintaining and managing mux cable. It currently runs linkmgrd binary that monitor and switches the mux if needed. This PR also introduces mux-container and starts linkmgrd as startup when build is configured with INCLUDE_MUX=y Edit: linkmgrd PR will follow. signed-off-by: Tamer Ahmed <tamer.ahmed@microsoft.com> Related work items: #2315, #3146150	2021-11-10 18:54:33 -08:00
tjchadaga	9a1b1bc44e	Fix for additional intf flap during fast-reboot (#9166 )	2021-11-09 23:20:06 +00:00
Lawrence Lee	8ada006302	[swss]: Start ndppd after vlanmgrd (#9155 ) Why I did it During swss container startup, if ndppd starts up before/with vlanmgrd, ndppd will be pinned at nearly 100% CPU usage. How I did it Only start ndppd after vlanmgrd is running. Also, call ndppd directly instead of through bash for improved logging and to prevent orphaned processes. Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2021-11-05 00:39:10 +00:00
Saikrishna Arcot	bb1bc59a22	docker-dhcp-relay: Fix waiting for interfaces to get set up (#9034 ) Fix the check used to wait for interfaces to come up. The group name in the supervisor config files has changed from isc-dhcp-relay to dhcp-relay. Also, in the wait script, wait 10 additional seconds after the vlans, port channels, and any interfaces are up. This is because dhcrelay listens on all interfaces (in addition to port channels and vlans), and to ensure that it stays in a clean state during runtime, wait some extra time to make sure that those interfaces are created as well. Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com>	2021-10-22 17:14:22 +00:00
kellyyeh	d4a6a009cf	Change radv interval to 3min (#8891 ) (cherry picked from commit `0e175e6d6c`)	2021-10-01 23:00:17 -07:00
kellyyeh	a4b6788b4b	Replace isc-dhcp with DHCPv6 Relay in dhcp_relay docker (#8884 )	2021-10-01 19:55:03 -07:00
kellyyeh	47ba7a9091	[dhcp_relay] DHCP relay support for IPv6 (#7772 ) (#8871 )	2021-09-30 01:33:02 -07:00
Christian Svensson	5dce093464	[mgmt-framework]: Fix typo in mgmt_vars.j2 (#8475 ) Signed-off-by: Christian Svensson <blue@cmd.nu>	2021-08-25 04:11:16 +00:00
Kostiantyn Yarovyi	387ae82c5d	[Pcied] run by python 3 Why I did it Pcied running by python 2. How I did it dropped python2 support and add python3 support for pcied in file docker-pmon.supervisord.conf.j2 How to verify it docker exec pmon supervisorctl status	2021-08-23 03:34:48 +00:00
xumia	b1c2659044	Support to build armhf/arm64 platforms on arm based system (#7731 ) (#8458 ) Why I did it Support to build armhf/arm64 platforms on arm based system without qemu simulator. When building the armhf/arm64 on arm based system, it is not necessary to use qemu simulator. How I did it Build armhf on armhf system, or build arm64 on arm64 system, by default, qemu simulator will not be used. When building armhf on arm64, and you have enabled armhf docker, then it will build images without simulator automatically. It is based how the docker service is run. Docker base image change: For amd64, change from debian:to amd64/debian: For arm64, change from multiarch/debian-debootstrap:arm64- to arm64v8/debian: For armhf, change from multiarch/debian-debootstrap:armhf- to arm32v7/debian: See https://github.com/docker-library/official-images#architectures-other-than-amd64 The mapping relations: arm32v6 --- armel arm32v7 --- armhf arm64v8 --- arm64 Docker image armhf deprecated info: https://hub.docker.com/r/armhf/debian, using arm32v7 instead.	2021-08-13 19:33:08 +08:00
richardyu	36ab000557	PTF adds unittest-xml-reporting (#8417 ) Co-authored-by: richardyu-ms <richard.yu@microsoft.com>	2021-08-12 07:09:58 +00:00
Sujin Kang	ae7fa32691	[pmon]: Enable Autorestart of the daemons in PMON for unexpected exit (#8358 ) Enable Autorestart of the daemons in PMON for unexpected exit Remove the daemon list from the critical_process which prevent the PMON from restarting when the individual daemon crashes.	2021-08-07 22:43:38 -07:00
Blueve	d2f2a07c7c	[ARM] Fix issue whre the ping6 tool is missing from orchagent docker (#8345 ) Signed-off-by: Jing Kan jika@microsoft.com	2021-08-05 15:25:53 +00:00
VenkatCisco	3aed7eab8f	[pmon]: add python3-jsonschema pmon (#8018 ) jsonschema is an implementation of JSON Schema for Python . Signed-off-by: Venkat Garigipati <venkatg@cisco.com>	2021-08-05 15:23:06 +00:00
novikauanton	08dc00f817	[iccpd][docker] fix initial startup configuration (#7982 ) #### Why I did it The process of config generation (sonic-cfggen) fails, but the services continue to run with invalid config #### How I did it * add exit with error on errors in start.sh script (because supervisord relies on start.sh return code). * fix jinja template. Jinja use common python expressions under the hood and `has_key` method was removed from dict in py3, so use check by `in` operator as it is supported by both py2 and py3. #### How to verify it * compile sonic with enabled iccp. * add mclag config to CONFIG_DB. ``` 'MC_LAG\|1' => { "local_ip": "10.0.0.2", "peer_ip": "10.0.0.3", "peer_link": "Ethernet8", "mclag_interface": "Ethernet12" } * unmaks, enable and start swss and iccpd services in sonic. * log in into the iccpd container and check the config file `/etc/iccpd/iccpd.conf` * expected config: ``` mclag_id:1 local_ip:10.0.0.2 peer_ip:10.0.0.3 peer_link:Ethernet8 mclag_interface:Ethernet12 system_mac:YOUR_SYSTEM_MAC #### Description for the changelog Fixed initial iccpd startup configuration.	2021-08-05 15:21:33 +00:00
Vivek Reddy	67202cc2bb	autorestart inside restapi docker is disabled (#8006 ) Fix issue with critical process in the restapi docker restarting immediately after getting killed Signed-off-by: Vivek Reddy Karri <vkarri@nvidia.com>	2021-07-27 05:14:28 +00:00
Guohan Lu	bed4c26b09	Revert "Add ethtool to docker-platform-monitor (#8017 )" This reverts commit `d66425dd76`.	2021-07-07 23:37:28 -07:00
VenkatCisco	d66425dd76	Add ethtool to docker-platform-monitor (#8017 ) #### Why I did it ethtool can be used to query and change settings such as speed, auto- negotiation and checksum offload on many network devices, especially Ethernet devices. #### How I did it add package extension to docker-platform-monitor/Dockerfile.j2	2021-07-07 09:40:11 +00:00
VenkatCisco	36d7dfbea3	Add libpci3 pkg to docker-platform-monitor (#8016 ) #### Why I did it The libpci library provides portable access to configuration registers of devices connected to the PCI bus. #### How I did it update dockers/docker-platform-monitor/Dockerfile.j2	2021-07-07 09:40:06 +00:00
thomas.cappleman@metaswitch.com	1d3e7ab161	[build]: Fix sonic-cfggen contextlib err (#7996 ) A recent version of contextlib2 (https://pypi.org/project/contextlib2/21.6.0/#history) has broken Python2 compatibility, so the version picked up by netaddr when using Python2 must be specified, or else builds fail Co-authored-by: Tom Zhu <tom.zhu@metaswitch.com>	2021-06-28 17:18:45 -07:00
Andriy Yurkiv	2fe91ae30f	Set default values only on the first start (#7735 )	2021-06-16 12:38:30 +00:00
bingwang-ms	c1b380df73	[docker-teamd]: Increase teammgrd timeout to allow graceful shutdown. (#7662 ) (#7842 ) The PR is a cherry-pick of #7662. Signed-off-by: Nazarii Hnydyn <nazariig@nvidia.com>	2021-06-10 12:49:18 -07:00
yozhao101	fb2c995f53	[202012][Monit] Deprecate the feature of monitoring the critical processes by Monit (#7823 ) Signed-off-by: Yong Zhao yozhao@microsoft.com Why I did it Currently we leveraged the Supervisor to monitor the running status of critical processes in each container and it is more reliable and flexible than doing the monitoring by Monit. So we removed the functionality of monitoring the critical processes by Monit. How I did it I removed the script process_checker and corresponding Monit configuration entries of critical processes. How to verify it I verified this on the device str-7260cx3-acs-1.	2021-06-09 09:04:22 -07:00
Myron Sosyak	e7009513da	[docker-database] Fix Python3 issue (#7700 ) #### Why I did it To avoid the following error ``` Traceback (most recent call last): File "/usr/local/bin/flush_unused_database", line 10, in <module> if 'PONG' in output: TypeError: a bytes-like object is required, not 'str' ``` `communicate` method returns the strings if streams were opened in text mode; otherwise, bytes. In our case text arg in Popen is not true and that means that `communicate` return the bytes #### How I did it Set `text=True` to get strings instead of bytes #### How to verify it run `/usr/local/bin/flush_unused_database` inside database container	2021-06-02 02:39:31 +00:00
bingwang-ms	eb8c05c306	Fix lldpmgrd syntax issue (#7742 ) Signed-off-by: bingwang <bingwang@microsoft.com>	2021-06-02 02:39:31 +00:00
Lawrence Lee	6a0e9078d4	[docker-orchagent]: Increase ndppd kernel poll interval (#7456 ) Why I did it ndppd by default reads /proc/net/ipv6_route ever 30 seconds. Since T1s advertise so many routes to ToRs, this file is extremely large, and reading it causes ndppd's CPU usage to spike every 30 seconds How I did it Increase the delay for reading this file to the maximum possible value (max integer value), which will result in CPU spikes every ~24 days instead of every 30 seconds How to verify it Start ndppd with the new config file, confirm that no CPU spikes are seen except at startup Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2021-06-02 02:38:54 +00:00
yozhao101	3af05fdffe	[Monit] Restart telemetry container if memory usage is beyond the threshold (#7645 ) Signed-off-by: Yong Zhao yozhao@microsoft.com Why I did it This PR aims to monitor the memory usage of streaming telemetry container and restart streaming telemetry container if memory usage is larger than the pre-defined threshold. How I did it I borrowed the system tool Monit to run a script memory_checker which will periodically check the memory usage of streaming telemetry container. If the memory usage of telemetry container is larger than the pre-defined threshold for 10 times during 20 cycles, then an alerting message will be written into syslog and at the same time Monit will run the script restart_service to restart the streaming telemetry container. How to verify it I verified this implementation on device str-7260cx3-acs-1.	2021-05-31 04:38:18 +00:00
bingwang-ms	c5d27750f2	Fix supervisor-proc-exit-listener startup issue in restapi (#7681 ) * Fix supervisor-proc-exit-listener startup issue in restapi Signed-off-by: bingwang <bingwang@microsoft.com>	2021-05-27 22:29:42 +00:00
LuiSzee	bc5b367d37	[radv] fix bug for radv can't startup if DEVICE_METADATA.localhost.type is NULL (#7651 ) Co-authored-by: Shi Lei <shil@centecnetworks.com>	2021-05-26 02:39:02 +00:00
Myron Sosyak	4ec13fd556	Fix python version (#7658 ) #### Why I did it To avoid the following logs ``` Mar 15 15:52:04.599302 igk-dut-04 INFO database#/supervisord: flushdb /bin/bash: /usr/local/bin/flush_unused_database: /usr/bin/python: bad interpreter: No such file or directory Mar 15 15:52:04.599947 igk-dut-04 INFO database#supervisord 2021-03-15 15:52:04,599 INFO exited: flushdb (exit status 126; not expected) ``` #### How I did it Fix shebang #### How to verify it Check the logs	2021-05-24 22:25:47 +00:00
xumia	2973a63f1d	Fix the type issue in rvtysh (#7648 ) Why I did it Change the type issue in the command rvtysh change PARA/para to PARAM/param	2021-05-24 22:25:47 +00:00
sudhanshukumar22	0a5551aabf	docker-lldp:intermittent DB errors will result in Client termination (#6119 ) This PR allows listen to hostname changes and mgmt ip changes.	2021-05-24 22:21:29 +00:00
abdosi	dbded1f48e	Changes in FRR temapltes for multi-asic (#6901 ) 1. Made the command next-hop-self force only applicable on back-end asic bgp. This is done so that BGPL iBGP session running on backend can send e-BGP learn nexthop. Back end asic FRR is able to recursively resolve the eBGP nexthop in its routing table since it knows about all the connected routes advertise from front end asic. 2. Made all front-end asic bgp use global loopback ip (Loopback0) as router id and back end asic bgp use Loopbacl4096 as ruter-id and originator id for Route-Reflector. This is done so that routes learnt by external peer do not see Loopback4096 as router id in show ip bgp <route-prerfix> output. 3. To handle above change need to pass Loopback4096 from BGP manager for jinja2 template generation. This was missing and this change/fix is needed for this also https://github.com/Azure/sonic-buildimage/blob/master/dockers/docker-fpm-frr/frr/bgpd/templates/dynamic/instance.conf.j2#L27 4. Enhancement to add mult_asic specific bgpd template generation unit test cases.	2021-05-24 21:59:57 +00:00
abdosi	8f6b3456ab	[multi-asic] BBR support on internal-peers for multi-asic platfroms. (#6848 ) Enable BBR config allowas-in 1 for internal peers Why I did: To advertise BBR routes learnt via e-BGP peer in one asic/namespace to another iBGP asic/namespace via Route Reflector.	2021-05-24 21:57:10 +00:00
VenkatCisco	91b4ce649e	[pmon]: add psmisc to bring fuser that dentifies processes that are using files or sockets (#7509 ) fuser support is required since new cisco hardware watchdog plugin uses them to check anyone else use's /dev/watchdogX resource. The actual validation happens in the platform code, but the package is required for pmon container. Currently the /dev/watchdogX is being used by cisco platform-monitor service. Cisco chassis level watchdog plugin uses "fuser" to claim the watchdog release from platform-monitor service.	2021-05-10 16:00:43 -07:00
Junchao-Mellanox	6e12c40f40	[Mellanox] Support new sensor conf file for MSN4700 A1/A0 (#7535 ) #### Why I did it MSN4700 A1/A0 used different sensor chip but keep the existing platform name x86_64-mlnx_msn4700-r0, this is a workaround to replace the sensor conf on MSN4700 A1/A0 #### How I did it Use a shell script to get the sensor conf path and copy that files to /etc/sensors.d/sensors.conf	2021-05-10 09:21:42 -07:00
trzhang-msft	d76206bae4	dhcpmon: support dual tor in docker template (#7470 )	2021-05-05 09:34:42 -07:00
judyjoseph	8cea931cad	Fixes for errors seen in staging devices (#7171 ) With the latest 201911 image, the following error was seen on staging devices with TSB command ( for both single asic, multi asic ). Though this err message doesn't affect the TSB functionality, it is good to fix. admin@STG01-0101-0102-01T1:~$ TSB BGP0 : % Could not find route-map entry TO_TIER0_V4 20 line 1: Failure to communicate[13] to zebra, line: no route-map TO_TIER0_V4 permit 20 % Could not find route-map entry TO_TIER0_V4 30 line 2: Failure to communicate[13] to zebra, line: no route-map TO_TIER0_V4 deny 30 In addition, in this PR I am fixing the message displayed to user when there are no BGP neighbors configured on that BGP instance. In multi-asic device there could be case where there are no BGP neighbors configured on a particular ASIC.	2021-05-03 13:19:29 -07:00
judyjoseph	7ae4a990e7	[docker-fpm-frr]: TSA/B/C changes for multi-asic (#6510 ) - Introduced TS common file in docker as well and moved common functions. - TSA/B/C scripts run only in BGP instances for front end ASICs. In addition skip enforcing it on route maps used between internal BGP sessions. admin@str--acs-1:~$ sudo /usr/bin/TSA System Mode: Normal -> Maintenance and in case of Multi-ASIC admin@str--acs-1:~$ sudo /usr/bin/TSA BGP0 : System Mode: Normal -> Maintenance BGP1 : System Mode: Normal -> Maintenance BGP2 : System Mode: Normal -> Maintenance	2021-05-03 13:19:17 -07:00
guxianghong	a0fde3a626	[arm] support compile sonic arm image on arm server (#7285 ) - Support compile sonic arm image on arm server. If arm image compiling is executed on arm server instead of using qemu mode on x86 server, compile time can be saved significantly. - Add kernel argument systemd.unified_cgroup_hierarchy=0 for upgrade systemd to version 247, according to #7228 - rename multiarch docker to sonic-slave-${distro}-march-${arch} Co-authored-by: Xianghong Gu <xgu@centecnetworks.com> Co-authored-by: Shi Lei <shil@centecnetworks.com>	2021-05-02 08:11:56 -07:00
xumia	1b05982727	Support readonly vtysh for sudoers (#7383 ) Why I did it Support readonly version of the command vtysh How I did it Check if the command starting with "show", and verify only contains single command in script.	2021-04-29 10:08:55 -07:00
kakkotetsu	e6bbb3c344	[restapi] fix python version during restapi startup (#7056 ) changed from python3 to python in supervisord.conf.	2021-04-22 14:36:09 -07:00
Vivek Reddy Karri	731401fe4f	Reverts the commit which reverts "Backport ethtool to support QSFP-DD (#5725 )" This reverts commit `a86cdd87cf`.	2021-04-15 19:15:58 -07:00
Stephen Sun	3cee45c298	[monit] Avoid monit error log by removing "-l" from monit_swss\|buffermgrd (#7236 ) Avoid the following error messages while dynamic buffer calculation is enabled ``` ERR monit[491]: 'swss\|buffermgrd' status failed (1) -- '/usr/bin/buffermgrd -l' is not running in host ``` Change /usr/bin/buffermgrd -l to /usr/bin/buffermgrd. The buffermgrd is started by -l for traditional model or -a for dynamic model. So we need to use the common section of both. Signed-off-by: Stephen Sun <stephens@nvidia.com>	2021-04-08 18:39:10 +00:00
Prince Sunny	e08dc12acf	[IPinIP] Add Loopback2 interface, change dscp mode to uniform (#7234 ) Co-authored-by: Ubuntu <prsunny>	2021-04-08 18:38:59 +00:00
Guohan Lu	a86cdd87cf	Revert "Backport ethtool to support QSFP-DD (#5725 )" This reverts commit `50e4cc1579`.	2021-04-01 13:11:15 -07:00
Joe LeVeque	dd9be59cd1	[202012][dockers][supervisor] Increase event buffer size for process exit listener; Set all event buffer sizes to 1024 (#7203 ) #### Why I did it Backport of https://github.com/Azure/sonic-buildimage/pull/7083 to the 202012 branch. To prevent error [messages](https://dev.azure.com/mssonic/build/_build/results?buildId=2254&view=logs&j=9a13fbcd-e92d-583c-2f89-d81f90cac1fd&t=739db6ba-1b35-5485-5697-de102068d650&l=802) like the following from being logged: ``` Mar 17 02:33:48.523153 vlab-01 INFO swss#supervisord 2021-03-17 02:33:48,518 ERRO pool supervisor-proc-exit-listener event buffer overflowed, discarding event 46 ``` This is basically an addendum to https://github.com/Azure/sonic-buildimage/pull/5247, which increased the event buffer size for dependent-startup. While supervisor-proc-exit-listener doesn't subscribe to as many events as dependent-startup, there is still a chance some containers (like swss, as in the example above) have enough processes running to cause an overflow of the default buffer size of 10. This is especially important for preventing erroneous log_analyzer failures in the sonic-mgmt repo regression tests, which have started occasionally causing PR check builds to fail. Example [here](https://dev.azure.com/mssonic/build/_build/results?buildId=2254&view=logs&j=9a13fbcd-e92d-583c-2f89-d81f90cac1fd&t=739db6ba-1b35-5485-5697-de102068d650&l=802). I set all supervisor-proc-exit-listener event buffer sizes to 1024, and also updated all dependent-startup event buffer sizes to 1024, as well, to keep things simple, unified, and allow headroom so that we will not need to adjust these values frequently, if at all.	2021-04-01 12:52:19 -07:00
Shi Su	576864b769	[bgp]: Reduce bgp connect retry timer to 10 seconds (#7169 ) The default bgp connect retry timer is 120 seconds. A reconnection will happen 120 seconds if the initial connection fails. This PR aims to allow a more frequent retry.	2021-03-31 08:49:23 -07:00
judyjoseph	902ad1357a	To decrease the Connect Retry Timer from default value which is 120sec to 10 sec. (#7087 ) Why I did it It was observed that on a multi-asic DUT bootup, the BGP internal sessions between ASIC's was taking more time to get ESTABLISHED than external BGP sessions. The internal sessions was coming up almost exactly 120 secs later. In multi-asic platform the bgp dockers ( which is per ASIC ) on switch start are bring brought up around the same time and they try to make the bgp sessions with neighbors (in peer ASIC's) which may be not be completely up. This results in BGP connect fail and the retry happens after 120sec which is the default Connect Retry Timer How I did it Add the command to set the bgp neighboring session retry timer to 10sec for internal bgp neighbors.	2021-03-26 17:39:50 +00:00
shlomibitton	50e4cc1579	Backport ethtool to support QSFP-DD (#5725 ) Backport ethtool debian package version 5.9 to support QSFP-DD cable parsing. Signed-off-by: Shlomi Bitton <shlomibi@nvidia.com>	2021-03-19 10:29:58 -07:00
trzhang-msft	1dec175743	[docker-dhcp-relay]: add -si support in dhcp docker template (#7053 )	2021-03-15 19:20:28 -07:00
Qi Luo	97426aff5a	[build]: Fix get-pip 2.7 url according to upstream announcement (#6999 ) ref: https://bootstrap.pypa.io/2.7/get-pip.py The URL you are using to fetch this script has changed, and this one will no longer work. Please use get-pip.py from the following URL instead: https://bootstrap.pypa.io/pip/2.7/get-pip.py	2021-03-10 09:32:12 -08:00
Tamer Ahmed	7ec9fbb678	Start DHCP Relay When Helpers IPs Are Available (#6961 ) #### Why I did it It is possible to have DHCP relay configuration with no servers/ helpers which result in DHCP container to crash. This PR fixes this issue by not starting DHCP relay for vlans with no DHCP helpers. resolves: #6931 closes: #6931 #### How I did it Do not add program group for dhcp relay with not dhcp helpers #### How to verify it Unit test	2021-03-10 09:23:56 -08:00
Qi Luo	414a942462	[radv] Disable radv for specific deployment_id (#6830 )	2021-02-23 23:56:01 +00:00
Andriy Yurkiv	569686ed84	Enable SAI_INGRESS_PRIORITY_GROUP_STAT_DROPPED_PACKETS counter by default (#6444 ) Signed-off-by: Andriy Yurkiv <ayurkiv@nvidia.com>	2021-02-23 23:56:01 +00:00
arlakshm	d7be5a021a	[Multi Asic] support of swss.rec and sairedis.rec for multi asic (#6310 ) Signed-off-by: Arvindsrinivasan Lakshmi Narasimhan arlakshm@microsoft.com - Why I did it This PR has the changes to support having different swss.rec and sairedis.rec for each asic. The logrotate script is updated as well - How I did it Update the orchagent.sh script to use the logfile name options in these PRs(Azure/sonic-swss#1546 and Azure/sonic-sairedis#747) In multi asic platforms the record files will be different for each asic, with the format swss.asic{x}.rec and sairedis.asic{x}.rec Update the logrotate script for multiasic platform .	2021-02-23 23:56:01 +00:00
Joe LeVeque	d7517a704c	[PDDF] Build and install Python 3 package (#6286 ) - Make PDDF code compliant with both Python 2 and Python 3 - Align code with PEP8 standards using autopep8 - Build and install both Python 2 and Python 3 PDDF packages	2021-02-23 23:56:01 +00:00
pra-moh	c240929607	[StreamingTelemetry] add noTLS support for debug purpose (#6704 ) adding noTLS mode for debugging purpose Removing config-set for port 8080. It fails to start telemetry if docker restarts in case on noTLS mode because it expects log_level config to be present as well.	2021-02-23 13:55:01 -08:00
yozhao101	cdef77f4c5	[SwSS] Disabled the autorestart of process `coppmgrd`. (#6774 ) coppmgrd process do not need to be auto-restarted if it exited unexpectedly. Signed-off-by: Yong Zhao <yozhao@microsoft.com>	2021-02-16 15:32:36 -08:00
Guohan Lu	89eb0c4a3c	[docker-fmp-frr]: remove blank lines in generated critical_process Signed-off-by: Guohan Lu <lguohan@gmail.com>	2021-01-28 09:29:09 -08:00
yozhao101	cc9c3f567e	[supervisord] Monitoring the critical processes with supervisord. (#6242 ) - Why I did it Initially, we used Monit to monitor critical processes in each container. If one of critical processes was not running or crashed due to some reasons, then Monit will write an alerting message into syslog periodically. If we add a new process in a container, the corresponding Monti configuration file will also need to update. It is a little hard for maintenance. Currently we employed event listener of Supervisod to do this monitoring. Since processes in each container are managed by Supervisord, we can only focus on the logic of monitoring. - How I did it We borrowed the event listener of Supervisord to monitor critical processes in containers. The event listener will take following steps if it was notified one of critical processes exited unexpectedly: The event listener will first check whether the auto-restart mechanism was enabled for this container or not. If auto-restart mechanism was enabled, event listener will kill the Supervisord process, which should cause the container to exit and subsequently get restarted. If auto-restart mechanism was not enabled for this contianer, the event listener will enter a loop which will first sleep 1 minute and then check whether the process is running. If yes, the event listener exits. If no, an alerting message will be written into syslog. - How to verify it First, we need checked whether the auto-restart mechanism of a container was enabled or not by running the command show feature status. If enabled, one critical process should be selected and killed manually, then we need check whether the container will be restarted or not. Second, we can disable the auto-restart mechanism if it was enabled at step 1 by running the commnad sudo config feature autorestart <container_name> disabled. Then one critical process should be selected and killed. After that, we will see the alerting message which will appear in the syslog every 1 minute. - Which release branch to backport (provide reason below if selected) 201811 201911 [x ] 202006	2021-01-28 09:28:27 -08:00
Shi Su	fc825f9a58	[FRR] Create a separate script to wait zebra to be ready to receive connections (#6519 ) The requirement for zebra to be ready to accept connections is a generic problem that is not specific to bgpd. Making the script to wait for zebra socket a separate script and let bgpd and staticd to wait for zebra socket.	2021-01-28 09:25:29 -08:00
Guohan Lu	0b2abafcde	[docker-ptf]: build docker ptf - combine docker-ptf-saithrift into docker-ptf docker - build docker-ptf under platform vs - remove docker-ptf for other platforms Signed-off-by: Guohan Lu <lguohan@gmail.com>	2021-01-28 09:23:12 -08:00
Tamer Ahmed	80ba3a4f6f	[dhcp-relay]: Launch DHCP Relay On L3 Vlan (#6527 ) Recent changes brought l2 vlan concept which do not have DHCP clients behind them and so DHCP relay is not required. Also, dhcpmon fails to launch on those vlans as their interfaces lack IP addresses. This PR limit launch of both DHCP relay and dhcpmon to L3 vlans only. singed-off-by: Tamer Ahmed <tamer.ahmed@microsoft.com>	2021-01-28 09:21:49 -08:00
Zhenhong Zhao	154a6ab6c5	[frrcfgd] introduce frrcfgd to manage frr config when frr_mgmt_framework_config is true (#5142 ) - Support for non-template based FRR configurations (BGP, route-map, OSPF, static route..etc) using config DB schema. - Support for save & restore - Jinja template based config-DB data read and apply to FRR during startup - How I did it - add frrcfgd service - when frr_mgmg_framework_config is set, frrcfgd starts in bgp container - when user changed the BGP or other related table entries in config DB, frrcfgd will run corresponding VTYSH commands to program on FRR. - add jinja template to generate FRR config file to be used by FRR daemons while bgp container restarted - How to verify it 1. Add/delete data on config DB and then run VTYSH "show running-config" command to check if FRR configuration changed. 1. Restart bgp container and check if generated FRR config file is correct and run VTYSH "show running-config" command to check if FRR configuration is consistent with attributes in config DB Co-authored-by: Zhenhong Zhao <zhenhong.zhao@dell.com>	2021-01-24 22:43:30 -08:00
Samuel Angebault	9dfee1f09a	[pmon]: Run ledd using python3 unless excluded (#6528 ) - Why I did it Ledd is the last daemon that is not enabled to run in python3. Even though there is a plan to deprecate this daemon and to replace it by something else it's one simple step toward python2 deprecation. - How I did it Changed the `command=` line for `ledd` in the `supervisord` configuration of `pmon`. Copied what was done for other daemons. - How to verify it Booting a product that has a `led_control.py` should now show the ledd running in python3. I ran `python3 -m pylint` on all `led_control.py` plugin which means that most of them should be python3 compliant. There is however still a risk that some might not work.	2021-01-22 11:42:22 -08:00
Shi Su	5079de7647	[bgpd]: Check zebra is ready to connect when starting bgpd (#6478 ) Fix #5026 There is a race condition between zebra server accepts connections and bgpd tries to connect. Bgpd has a chance to try to connect before zebra is ready. In this scenario, bgpd will try again after 10 seconds and operate as normal within these 10 seconds. As a consequence, whatever bgpd tries to sent to zebra will be missing in the 10 seconds. To avoid such a scenario, bgpd should start after zebra is ready to accept connections.	2021-01-19 01:11:50 -08:00
pavel-shirshov	cd8417afd7	[docker-frr]: Use egrep with regexp to match correct TSA rules (#6403 ) - Why I did it Earlier today we found a bug in the SONiC TSA implementation. TSC shows incorrect output (see below) in case we have a route-map which contains TSA route-map as a prefix. ``` admin@str-s6100-acs-1:~$ TSC Traffic Shift Check: System Mode: Not consistent ``` The reason is that TSC implementation has too loose regexps in TSA utilities, which match wrong route-map entries: For example, current TSC matches following ``` route-map TO_BGP_PEER_V4 permit 200 route-map TO_BGP_PEER_V6 permit 200 ``` But it should match only ``` route-map TO_BGP_PEER_V4 permit 20 route-map TO_BGP_PEER_V4 deny 30 route-map TO_BGP_PEER_V6 permit 20 route-map TO_BGP_PEER_V6 deny 30 ``` - How I did it I fixed it by using egrep with `^` and `$` regexp markers which match begin and end of the line. - How to verify it 1. Add follwing entry to FRR config: ``` str-s6100-acs-1# str-s6100-acs-1# conf t str-s6100-acs-1(config)# route-map TO_BGP_PEER_V4 permit 200 str-s6100-acs-1(config-route-map)# end ``` 2. Use the TSC command and check output. It should show normal. ``` admin@str-s6100-acs-1:~$ TSC Traffic Shift Check: System Mode: Normal```	2021-01-15 08:20:14 -08:00
carl-nokia	d2f684b05c	[Platform][nokia]: python3-smbus package add with python3 and jinja fixes (#6416 ) fix platform driver breakage due to python3 upgrade and fix load minigraph errors with config load_minigraph -y - How I did it added python3-smbus to the pmon docker template since the previous was python2 specific fixed additional "ord" python2 specific code fixed the jinja templates used by qos reload - the template logic required data to be parsed - How to verify it run "show platform XXX" commands and verify output run "sudo config load_minigraph -y" and verify configuration run "show interfaces XXX" and verify output Co-authored-by: Carl Keene <keene@nokia.com>	2021-01-15 08:16:32 -08:00
pavel-shirshov	03391f20c5	[bgpcfgd]: Support default action for "Allow prefix" feature (#6370 ) * Use 20 and 30 route-map entries instead of 2 and 3 for TSA * Added support for dynamic "Allow list" default action. Co-authored-by: Pavel Shirshov <pavel.contrib@gmail.com>	2021-01-09 08:29:19 -08:00
abdosi	be82cdbad5	Updated imfile configuration for supervisord logs (#6368 ) Updated imfile configuration for supervisord logs for stretch and buster.	2021-01-09 08:27:08 -08:00
sudhanshukumar22	c111a68e74	[docker-lldp]: sonic advertise meaningful SysDescription instead of debian (#6114 ) Sonic devices advertise meaningful system description along with Debian package information. before the fix: ------------- admin@sonic:~$ show lldp neighbors ------------------------------------------------------------------------------- LLDP neighbors: ------------------------------------------------------------------------------- Interface: Ethernet0, via: LLDP, RID: 3, Time: 0 day, 16:36:30 SysName: sonic SysDescr: Debian GNU/Linux 9 (stretch) Linux 4.9.0-11-2-amd64 #1 SMP Debian 4.9.189-3+deb9u2 (2019-11-11) x86_64 ------------------------------------------------------------------------------- After the fix: root@sonic:~# show lldp neighbors Ethernet16 ------------------------------------------------------------------------------- LLDP neighbors: ------------------------------------------------------------------------------- Interface: Ethernet16, via: LLDP, RID: 10, Time: 0 day, 00:01:00 SysName: sonic SysDescr: SONiC Software Version: SONiC.sonic_upstream_1.0_daily_201130_1501_62-dirty-20201130.203529 - HwSku: Accton-AS7816-64X - Distribution: Debian 10.6 - Kernel: 4.19.0-9-2-amd64 ------------------------------------------------------------------------------- Signed-off-by: sudhanshukumar22 <sudhanshu.kumar@broadcom.com>	2021-01-09 08:26:21 -08:00
abdosi	dd25f774b5	[rsyslog]: Explicitly set the notify mode for rsyslog imfile module (#6351 ) Enable the notify mode of rsyslogd imfile module used for supervisord logs in docker container. Setup the mode="inotify" when loading imfile, made sure we are are getting supervisord logs in host immediately. Signed-off-by: Abhishek Dosi <abdosi@microsoft.com>	2021-01-06 06:19:50 -08:00
abdosi	ef0088c29f	Enable the notify mode of rsyslogd imfile module used for supervisord (#6298 ) Enable the notify mode of rsyslogd imfile module used for supervisord logs in docker container	2020-12-31 17:01:57 -08:00
Ubuntu	273846a412	FRR 7.5 Build libyang1 which is required for frr 7.5	2020-12-29 03:44:49 -08:00
Stepan Blyshchak	23f1d51de3	[ipinip.json.j2] align mellanox configuration dst_ip with other platforms (#6304 ) Mellanox already supports multiple destination IPs in IPinIP tunnel configuration, thus removing mellanox exception for IPinIP configuration. - How I did it Removed "dst_ip" field generation in mellanox platform condition. Sorted the "dst_ip" list, so that it is easier to test against sample configuration in unit tests. Aligned unit test sample. Signed-off-by: Stepan Blyschak <stepanb@nvidia.com>	2020-12-28 20:53:12 -08:00
Travis Van Duyn	6efc0a885f	Convert snmp.yml to configdb (#6205 ) This PR is in preparation to move from snmp.yml to configdb. This will more closely align with other commands in sonic and use configdb as the source of truth for snmp configuration. Note: This is the first of 2 PR's to enable this. This PR will not change any functionality but will allow the snmp.yml file info to be put into the configdb. Created a script that takes the snmp.yml variables and converts them to the configdb format. Added file to dockerfile.j2 so that file is copied in the container. Updated start.sh file to automatically run the python conversion script each time the docker container is restarted.	2020-12-28 11:51:58 -08:00
Guohan Lu	ed58684e36	[docker-frr]: add static ipv6 loopback route to allow bgp to advertise prefix frr does not advertise route if local route is not reachable, as a result loopback route /64 is not advertised to the neighbors. Add static route allows frr to advertise the route to its peers Signed-off-by: Guohan Lu <lguohan@gmail.com>	2020-12-28 10:34:34 -08:00
Junchao-Mellanox	51f896b33e	Add pmon daemons python3 build support (#6176 ) - Why I did it python2 is end of life and SONiC is going to support python3. This PR is going to support: 1. Build pmon daemons with python3 2. Install and run python3 version pmon daemons - How I did it 1. Change pmon daemons make files to build bothe python2 and python3 whl 2. Change docker-platform-monitor make files to install both python2 and python3 whl 3. Change pmon docker startup files to start pmon daemons according to the supported platform API version	2020-12-28 10:19:24 -08:00
Prince Sunny	8fd50e895c	[submodule]: swss Tunnel Manager changes (#5843 ) Introduce tunnel manager daemon. Start the process as part of swss container Submodule update for swss: 9ed3026 - 2020-12-24 : [NAT] ACL Rule with DO_NOT_NAT action is getting failed. (#1502) [Akhilesh Samineni] c39a4b1 - 2020-12-23 : Mux/IPTunnel orchagent changes (#1497) [Prince Sunny] bc8df0e - 2020-12-23 : Add support for headroom pool watermark (#1567) [Neetha John]	2020-12-26 11:17:18 -08:00
Joe LeVeque	d40c9a1e8d	[docker-base-buster][docker-config-engine-buster] No longer install Python 2 (#6162 ) - Why I did it As part of migrating SONiC codebase from Python 2 to Python 3 - How I did it - No longer install Python 2 in docker-base-buster or docker-config-engine-buster. - Install Python 2 and pip2 in the following containers until we can completely eliminate it there: - docker-platform-monitor - docker-sonic-mgmt-framework - docker-sonic-vs - Pin pip2 version <21 where it is still temporarily needed, as pip version 21 will drop support for Python 2 - Also preform some other cleanup, ensuring that pip3, setuptools and wheel packages are installed in docker-base-buster, and then removing any attempts to re-install them in derived containers	2020-12-25 21:29:25 -08:00
KISHORE KUNAL	4bb8ab3495	Add support to start fdbsyncd when orchagent docker starts (#5979 ) Add support to start fdbsyncd when swss docker starts. New demon is added to sync MAC from Kernel to DB and vise versa.	2020-12-24 18:36:01 -08:00
Wei Bai	8939202f67	[docker-sonic-mgmt]: Upgrade Tgen API from 0.0.42 to 0.0.70 (#6275 ) Tgen API 0.0.42 has many problems. We have fixed them in 0.0.70.	2020-12-24 01:53:31 -08:00
faraazbrcm	9d35fa19dc	[mgmt-framework]: support python3 in mgmt-framework (#6038 ) Fix specific version for mmh3 for python2 and python3 and Add pyang for python3	2020-12-22 12:06:28 -08:00
Renuka Manavalan	ba02209141	First cut image update for kubernetes support. (#5421 ) * First cut image update for kubernetes support. With this, 1) dockers dhcp_relay, lldp, pmon, radv, snmp, telemetry are enabled for kube management init_cfg.json configure set_owner as kube for these 2) Each docker's start.sh updated to call container_startup.py to register going up As part of this call, it registers the current owner as local/kube and its version The images are built with its version ingrained into image during build 3) Update all docker's bash script to call 'container start/stop/wait' instead of 'docker start/stop/wait'. For all locally managed containers, it calls docker commands, hence no change for locally managed. 4) Introduced a new ctrmgrd service, that helps with transition between owners as kube & local and carry over any labels update from STATE-DB to API server 5) hostcfgd updated to handle owner change 6) Reboot scripts are updatd to tag kube running images as local, so upon reboot they run the same image. 7) Added kube_commands.py to handle all updates with Kubernetes API serrver -- dedicated for k8s interaction only.	2020-12-22 08:01:33 -08:00
xumia	0a36de3a89	Recover "Support SONiC Reproduceable Build-debian/pip/web packages (#6255 ) * Revert "Revert "Support SONiC Reproduceable Build-debian/pip/web packages (#5718)"" This reverts commit `17497a65e3`. * Revert "Revert "Remove unnecessary sudo authority in build Makefile (#6237)"" This reverts commit `163b7111b5`.	2020-12-21 15:31:10 +08:00
Guohan Lu	17497a65e3	Revert "Support SONiC Reproduceable Build-debian/pip/web packages (#5718 )" This reverts commit `55a707586b`.	2020-12-18 23:37:27 -08:00
macikgozwa	169b2fb188	[docker-ptf]: Updating Python-based GNMI client (#6216 ) Upgrading the reference for the Python GNMI tool repository. The commit for the new payload Co-authored-by: Murat Acikgoz <muacikgo@microsoft.com>	2020-12-17 22:12:04 -08:00
xumia	55a707586b	Support SONiC Reproduceable Build-debian/pip/web packages (#5718 ) * Support SONiC reproduceable build for deb/py2/py3/web * Remove j2 files * Fix bug * Fix some issues 1. Change some code format issues 2. Fix curl calling wget command, pip2 calling pip3 issue 3. Fix wget/curl downloading multiple urls issue * Fix some code format issue * Fix bug * Fix bug * Fix command path hard code in build info scripts issue * Add debian package sonic-build-tools * Fix auto debian package removed issue * Change build debian package name, and change the folder * Collect the pre-versions and post-versions * Change to use debian:buster * Remove apt-mark and improve code * Remove set_build_hooks * Change docker trusted gpg files * Fix docker build COPY directory name issue * Move the trusted gpg files into the sonic-build-hooks package	2020-12-17 13:06:53 +08:00
mprabhu-nokia	41012f791e	In modular chassis, add CHASSIS_STATE_DB on control card (#5624 ) HLD: Azure/SONiC#646 In modular chassis, add CHASSIS_STATE_DB on control card Why I did it Modular Chassis has control-cards, line-cards and fabric-cards along with other peripherals. Control-Card CHASSIS_STATE_DB will be the central DB to maintain any state information of cards that is accessible to control-card/ How I did it Adding another DB on an existing REDIS instance running on port 6380.	2020-12-15 17:15:00 -08:00
mprabhu-nokia	00cea080af	Chassisd to monitor cards in a modular chassis (#5523 ) HLD: Azure/SONiC#646 Introducing chassisd process to monitor status of the control, line and fabric cards in a modular chassis. - Why I did it Modular Chassis has control-cards, line-cards and fabric-cards along with other peripherals. Chassisd will be a central entity that has visibility of the entire chassis. - How I did it Chassisd process will monitor cards in the main thread. Another configuation_handling_task is created to listen to CONFIG_DB for admin_status up/down events. The monitored status is persisted in REDIS-DB.	2020-12-15 16:28:58 -08:00
zhenggen-xu	182a809dc3	[docker-vs][docker-orchagent] install python3 dependent packages for restore_neighbors.py (#6207 ) Install the necessary python3 dependent packages to convert restore_neighbor.py to support python3 as python2 is EOL. See: Azure/sonic-swss#1542 Signed-off-by: Zhenggen Xu <zxu@linkedin.com>	2020-12-15 11:06:30 -08:00
Sabareesh-Kumar-Anandan	9f4ca01388	[sonic-config-engine] Adding dependent pkgs needed for arm compilation (#6186 ) libxslt-dev and libz-dev are dependencies for lxml==4.6.1 which is required for pyangbind==0.8.1 lxml-4.6.2-cp37-cp37m-manylinux1_x86_64.whl is directly downloaded in amd64 whereas in arm this is built from lxml-4.6.2.tar.gz Signed-off-by: Sabareesh Kumar Anandan <sanandan@marvell.com>	2020-12-15 08:44:46 -08:00
Stephen Sun	e010d83fc3	[Dynamic buffer calc] Support dynamic buffer calculation (#6194 ) - Why I did it To support dynamic buffer calculation. This PR also depends on the following PRs for sub modules - [sonic-swss: [buffermgr/bufferorch] Support dynamic buffer calculation #1338](https://github.com/Azure/sonic-swss/pull/1338) - [sonic-swss-common: Dynamic buffer calculation #361](https://github.com/Azure/sonic-swss-common/pull/361) - [sonic-utilities: Support dynamic buffer calculation #973](https://github.com/Azure/sonic-utilities/pull/973) - How I did it 1. Introduce field `buffer_model` in `DEVICE_METADATA\|localhost` to represent which buffer model is running in the system currently: - `dynamic` for the dynamic buffer calculation model - `traditional` for the traditional model in which the `pg_profile_lookup.ini` is used 2. Add the tables required for the feature: - ASIC_TABLE in platform/\<vendor\>/asic_table.j2 - PERIPHERAL_TABLE in platform/\<vendor\>/peripheral_table.j2 - PORT_PERIPHERAL_TABLE on a per-platform basis in device/\<vendor\>/\<platform\>/port_peripheral_config.j2 for each platform with gearbox installed. - DEFAULT_LOSSLESS_BUFFER_PARAMETER and LOSSLESS_TRAFFIC_PATTERN in files/build_templates/buffers_config.j2 - Add lossless PGs (3-4) for each port in files/build_templates/buffers_config.j2 3. Copy the newly introduced j2 files into the image and rendering them when the system starts 4. Update the CLI options for buffermgrd so that it can start with dynamic mode 5. Fetches the ASIC vendor name in orchagent: - fetch the vendor name when creates the docker and pass it as a docker environment variable - `buffermgrd` can use this passed-in variable 6. Clear buffer related tables from STATE_DB when swss docker starts 7. Update the src/sonic-config-engine/tests/sample_output/buffers-dell6100.json according to the buffer_config.j2 8. Remove buffer pool sizes for ingress pools and egress_lossy_pool Update the buffer settings for dynamic buffer calculation	2020-12-13 11:35:39 -08:00
Dong Zhang	b2a3de5f4f	[MultiDB] add mutidb warmboot support - restoring database (#5773 ) * restoring each database with all data before warmboot and then flush unused data in each instance, following the multiDB warmboot design at https://github.com/Azure/SONiC/blob/master/doc/database/multi_database_instances.md * restore needs to be done in database docker since we need to know the database_config.json in new version * copy all data rdb file into each instance restoration location andthen flush unused database * other logic is the same as before * backing up database part is in another PR at sonic-utilities https://github.com/Azure/sonic-utilities/pull/1205, they depend on each other	2020-12-10 11:06:19 -08:00
trzhang-msft	d4d90a8963	Support for dual tor option in dhcp docker template (#6152 )	2020-12-09 18:10:00 -08:00
Samuel Angebault	8576911a57	[database-chassis]: Fix the way database-chassis start (#6099 ) The service crash when the platform boots due to missing waits. /usr/bin/database.sh tries to operate on a missing socket and fails. We now wait for the chassis database to be ready the same way we do database.	2020-12-04 10:09:35 -08:00
Joe LeVeque	83f0d8240e	[pmon]: Install vanilla 'thrift' Python 2 and 3 packages for Barefoot in host and PMon (#6080 ) Barefoot platform vendors' sonic_platform packages import the Python 'thrift' library. Previously, our custom-built package was being installed in the PMon container and host OS. However, we are only building a Python 2 version of that package, which was only intended for use with saithrift. Fixes #6077	2020-12-04 08:41:17 -08:00
Joe LeVeque	905a5127bb	[Python] Align files in root dir, dockers/ and files/ with PEP8 standards (#6109 ) - Why I did it Align style with slightly modified PEP8 standards (extend maximum line length to 120 chars). This will also help in the transition to Python 3, where it is more strict about whitespace, plus it helps unify style among the SONiC codebase. Will tackle other directories in separate PRs. - How I did it Using `autopep8 --in-place --max-line-length 120` and some manual tweaks.	2020-12-03 15:57:50 -08:00

1 2 3 4 5 ...

877 Commits