sonic-buildimage

Author	SHA1	Message	Date
Ying Xie	64ce6696bb	[mux] skip mux operations during warm shutdown (#11937 ) * [mux] skip mux operations during warm shutdown - Enhance write_standby.py script to skip actions during warm shutdown. - Expand the support to BGP service. - MuX support was added by a previous PR. - don't skip action during warm recovery Signed-off-by: Ying Xie <ying.xie@microsoft.com>	2022-10-03 22:30:55 +00:00
Longxiang Lyu	893391f76e	[mux] Exit to write `standby` state to `active-active` ports (#11821 ) [mux] Exit to write standby state to `active-active` ports Signed-off-by: Longxiang Lyu <lolv@microsoft.com>	2022-10-03 19:52:51 +00:00
mssonicbld	a05d86a047	[ci/build]: Upgrade SONiC package versions (#12252 )	2022-10-02 18:38:46 +08:00
mssonicbld	f61c84b09e	[ci/build]: Upgrade SONiC package versions (#12243 )	2022-09-30 21:36:39 +08:00
mssonicbld	8994298788	[ci/build]: Upgrade SONiC package versions (#12199 )	2022-09-28 21:32:50 +08:00
mssonicbld	dc67ec2869	[ci/build]: Upgrade SONiC package versions (#12181 )	2022-09-25 20:10:04 +08:00
zitingguo-ms	95b19bbb46	Pick up fixes and make up BRCM SAI version to 4.3.7.1 (#12069 ) Signed-off-by: zitingguo-ms <zitingguo@microsoft.com> Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-09-22 22:59:51 -07:00
mssonicbld	2b5eb45d83	[ci/build]: Upgrade SONiC package versions (#12141 )	2022-09-21 20:37:27 +08:00
mssonicbld	4396d67d64	[ci/build]: Upgrade SONiC package versions (#12120 )	2022-09-20 19:51:45 +08:00
mssonicbld	6678e3356e	[ci/build]: Upgrade SONiC package versions (#12101 )	2022-09-18 19:48:13 +08:00
lixiaoyuner	0abf8d0419	Make client indentity by AME cert (#11946 ) * Make client indentity by AME cert * Join k8s cluster by ipv6 * Change join test cases * Test case bug fix * Improve read node label func * Configure kubelet and change test cases * For kubernetes version 1.22.2 * Fix undefine issue Signed-off-by: Yun Li <yunli1@microsoft.com>	2022-09-17 00:41:53 +00:00
mssonicbld	a195f8343f	[ci/build]: Upgrade SONiC package versions (#12068 )	2022-09-14 23:05:43 +08:00
mssonicbld	37ada92ad5	[ci/build]: Upgrade SONiC package versions (#12043 )	2022-09-11 18:59:58 +08:00
mssonicbld	5e8332f505	[ci/build]: Upgrade SONiC package versions (#12039 )	2022-09-09 21:46:28 +08:00
siqbal1986	10b6a5c402	[202012] StateDB table cleanup for VNET routes (#11999 ) * added cleanup of tables. 'VNET_ROUTE_TUNNEL_TABLE', 'VNET_ROUTE_TABLE'	2022-09-08 14:57:20 -07:00
mssonicbld	143af80061	[ci/build]: Upgrade SONiC package versions (#11976 )	2022-09-07 00:40:48 +08:00
mssonicbld	586a623422	[ci/build]: Upgrade SONiC package versions (#11911 )	2022-09-04 19:29:06 +08:00
Lawrence Lee	e821dd8551	[arp_update]: Set failed IPv6 neighbors to incomplete (#11919 ) After pinging any failed IPv6 neighbor entries, set the remaining failed/incomplete entries to a permanent INCOMPLETE state. This manual setting to INCOMPLETE prevents these entries from automatically transitioning to FAILED state, and since they are now incomplete any subsequent NA messages for these neighbors is able to resolve the entry in the cache. Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2022-09-02 21:57:47 +00:00
Ying Xie	4ab83170a5	[write_standby] update write_standby.py script (#11650 ) Why I did it The initial value has to be present for the state machines to work. In active-standby dual-tor scenario, or any hardware mux scenario, the value will be updtaed eventually with a delay. However, in active-active dual-tor scenario, there is no other mechanism to initialize the value and get state machines started. So this script will have to write something at start up time. For active-active dualtor, 'active' is a more preferred initial value, the state machine will switch the state to standby soon if link prober found link not in good state. How I did it Update the script to always provide initial values. How to verify it Tested on active-active dual-tor testbed. Signed-off-by: Ying Xie ying.xie@microsoft.com	2022-09-01 23:57:23 +00:00
Jing Zhang	9d3194c77a	Avoid write_standby in warm restart context (#11283 ) Avoid write_standby in warm restart context. sign-off: Jing Zhang zhangjing@microsoft.com Why I did it In warm restart context, we should avoid mux state change. How I did it Check warm restart flag before applying changes to app db. How to verify it Ran write_standby in table missing, key missing, field missing scenarios. Did a warm restart, app db changes were skipped. Saw this in syslog: WARNING write_standby: Taking no action due to ongoing warmrestart.	2022-09-01 23:57:17 +00:00
mssonicbld	ed68e4c97c	[ci/build]: Upgrade SONiC package versions (#11896 )	2022-08-30 22:44:47 +08:00
mssonicbld	347b2dddcd	[ci/build]: Upgrade SONiC package versions (#11757 )	2022-08-29 14:08:14 +08:00
mssonicbld	07082bb5f5	[ci/build]: Upgrade SONiC package versions (#11676 )	2022-08-16 13:07:32 +00:00
Stepan Blyshchak	8ab448a852	[swss.sh/syncd.sh] Trap only on EXIT (#11590 ) When using trap on SIGTERM the script will not react to the SIGTERM signal sent while a child is executing. I.e, the following script does not react on SIGTERM sent to it if it is waiting for sleep to finish: ``` trap "echo Handled SIGTERM" 0 2 3 15 echo "Before sleep" sleep inf echo "After sleep" ``` Instead, trap only on EXIT which covers also a scenario with exit on SIGINT, SIGTERM. Signed-off-by: Stepan Blyschak <stepanb@nvidia.com>	2022-08-11 20:38:20 +00:00
zitingguo-ms	5b5bd5e818	[202012 BRCM SAI 4.3.7.0] Pick up fixes and make up BRCM SAI version to 4.3.7.0 (#11681 ) Pick upfollowing fixes and update BRCM SAI to 4.3.7.0: CS00012208537: Add back previous commit 54c5bc4848eb748 CS00012253061,SONIC-63280: WB from 3.5 to 4.3, followed by WB to 4.3 CS00012207978: SDK-296517, time spent for SAI operations CS00012245601,SONIC-62898: Egress ACL Counted ad Interface TX drops Update pcbb with Fixes for CS00012243699 Upgrade on pcbb with Fixes for KB0025353, CS00012221689, CS00012221688, KB0025391, CS00012230519 commit of "CS00012221688:PFC frames egressing, PFC storm happens simultaneously on 2 ports" is purposely skipped to be picked up later due to SWSS dependency not ready. Why I did it How I did it How to verify it Tested build target, successful Manually run these tests after installing sai binary within image 20201231.73 on 7050CX3 (TD3) T0 DUT, all passed. vxlan/test_vxlan_decap.py fdb/test_fdb.py pfcwd/test_pfcwd_all_port_storm.py acl/null_route/test_null_route_helper.py acl/test_acl.py vlan/test_vlan.py platform_tests/test_reboot.py Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-08-10 15:02:47 -07:00
Jing Zhang	ffd9e190e1	Update WARM START FINALIZER to wait for linkmgrd to reconcile (#11477 ) Spanning from sonic-net/sonic-linkmgrd#76, this PR is to update warm restart finalizer to wait for linkmgrd to be reconciled. sign-off: Jing Zhang zhangjing@microsoft.com Why I did it To make sure finalizer save config after linkmgrd's reconciliation. How I did it Add linkmgrd to the reconciliation wait list of warmboot finalizer. How to verify it Verified on lab device, linkmgrd reconciled as expected.	2022-08-09 21:05:12 +00:00
mssonicbld	14f93e15c6	[ci/build]: Upgrade SONiC package versions (#11629 ) Why I did it Upgrade SONiC Versions	2022-08-07 11:27:16 +08:00
Lawrence Lee	04ba6da1ab	[202012][arp_update]: Resolve failed neighbors on dualtor (#11641 ) In arp_update, check for FAILED or INCOMPLETE kernel neighbor entries and manually ping them to try and resolve the neighbor Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2022-08-05 23:30:04 -07:00
tjchadaga	6d66d9b8fc	Revert "Add load_minigraph option to include traffic-shift-away during config migration (#11403 )" (#11625 ) This reverts commit `6c2f99a327`.	2022-08-06 10:05:45 +05:30
bingwang-ms	84aca00847	[202012]Support different `DSCP_TO_TC_MAP` for T1 in dualtor deployment (#11580 ) Why I did it This PR is to backport #11569 into 202012 branch. This PR is to apply different DSCP_TO_TC_MAP to downlink and uplink ports on T1 in dualtor deployment. For T1 downlink ports (To T0) The DSCP_TO_TC_MAP is not changed. DSCP2 and DSCP6 are mapped to TC2 and TC6 respectively. For T1 uplink ports (To T1) A new DSCP_TO_TC_MAP\|AZURE_UPLINK is defined and applied. DSCP2 and DSCP6 are mapped to TC1 to avoid mixing up lossy and lossless traffic from T2. The extra lossy PG2 and PG6 added in PR #11157 is reverted as well because no traffic from T2 is mapped to PG2 or PG6 now. How I did it Define a new map DSCP_TO_TC_MAP\|AZURE_UPLINK for 7260 T1. How to verify it Verified by test case in test_j2files.py.	2022-08-01 08:59:45 -07:00
Nikola Dancejic	c5a5734242	[swss] Adding bgp container as dependent of swss (#11168 ) What I did: Added bgp as a dependent of swss Why I did it: bgp container was not restarting on swss crash. When swss crashes, linkmgrd doesn't initate a switchover because it cannot access the default route from orchagent. Bringing down bgp with swss will isolate the ToR, causing linkmgrd to initiate a switchover to the peer ToR avoiding significant packet loss. Signed-off-by: Nikola Dancejic <ndancejic@microsoft.com>	2022-07-29 09:37:09 -07:00
Lior Avramov	a40aca43b9	[memory_checker] Do not check memory usage of containers if docker daemon is not running (#11476 ) Fix in Monit memory_checker plugin. Skip fetching running containers if docker engine is down (can happen in deinit). This PR fixes issue #11472. Signed-off-by: liora liora@nvidia.com Why I did it In the case where Monit runs during deinit flow, memory_checker plugin is fetching the running containers without checking if Docker service is still running. I added this check. How I did it Use systemctl is-active to check if Docker engine is still running. How to verify it Use systemctl to stop docker engine and reload Monit, no errors in log and relevant print appears in log. Which release branch to backport (provide reason below if selected) The fix is required in 202205 and 202012 since the PR that introduced the issue was cherry picked to those branches (#11129).	2022-07-27 23:28:19 +00:00
tjchadaga	6c2f99a327	Add load_minigraph option to include traffic-shift-away during config migration (#11403 )	2022-07-27 23:27:21 +00:00
bingwang-ms	c5eb031111	[202012] Add flag to control the generation of global level map (#11451 ) Why I did it This PR is to cherry-pick #11448 to 202012 branch after resolving conflicts. There are conflicts in files/build_templates/qos_config.j2 src/sonic-config-engine/tests/test_j2files.py	2022-07-15 09:44:45 -07:00
mssonicbld	550ab26fc7	[ci/build]: Upgrade SONiC package versions (#11422 )	2022-07-12 15:39:32 +00:00
Neetha John	26ee4ae4a4	Add backend acl template (#11220 ) Why I did it Storage backend has all vlan members tagged. If untagged packets are received on those links, they are accounted as RX_DROPS which can lead to false alarms in monitoring tools. Using this acl to hide these drops. How I did it Created a acl template which will be loaded during minigraph load for backend. This template will allow tagged vlan packets and dropped untagged How to verify it Unit tests Signed-off-by: Neetha John <nejo@microsoft.com>	2022-07-08 21:39:39 +00:00
mssonicbld	9a86fa9264	[ci/build]: Upgrade SONiC package versions (#11074 ) Upgrade SONiC Versions	2022-07-06 11:00:50 +08:00
xumia	32cda89f93	[Build]: Support to use symbol links for lazy installation targets to reduce the image size (#10923 ) Why I did it Support to use symbol links in platform folder to reduce the image size. The current solution is to copy each lazy installation targets (xxx.deb files) to each of the folders in the platform folder. The size will keep growing when more and more packages added in the platform folder. For cisco-8000 as an example, the size will be up to 2G, while most of them are duplicate packages in the platform folder. How I did it Create a new folder in platform/common, all the deb packages are copied to the folder, any other folders where use the packages are the symbol links to the common folder. Why platform.tar? We have implemented a patch for it, see #10775, but the problem is the the onie use really old unzip version, cannot support the symbol links. The current solution is similar to the PR 10775, but make the platform folder into a tar package, which can be supported by onie. During the installation, the package.tar will be extracted to the original folder and removed.	2022-07-05 20:57:49 +00:00
yozhao101	4487a962e3	[memory_checker] Do not check memory usage of containers which are not created (#11129 ) Signed-off-by: Yong Zhao yozhao@microsoft.com Why I did it This PR aims to fix an issue (#10088) by enhancing the script memory_checker. Specifically, if container is not created successfully during device is booted/rebooted, then memory_checker do not need check its memory usage. How I did it In the script memory_checker, a function is added to get names of running containers. If the specified container name is not in current running container list, then this script will exit without checking its memory usage. How to verify it I tested on a lab device by following the steps: Stops telemetry container with command sudo systemctl stop telemetry.service Removes telemetry container with command docker rm telemetry Checks whether the script memory_checker ran by Monit will generate the syslog message saying it will exit without checking memory usage of telemetry.	2022-07-05 20:57:45 +00:00
Samuel Angebault	d15a484dfa	[202012][Arista] Fix cmdline generation during warm-reboot from 201811/201911 (#11161 ) Issue fixed: when performing a warm-reboot or fast-reboot from 201811 or 201911 to 202012 the kernel command line contains duplicate information. This issue is related to a change that was made to make 202012 boot0 file more futureproof. A cold reboot brings everything back into a clean slate though not always desirable. Changes done: Added some logic to properly detect the end of the Aboot cmdline when cmdline-aboot-end delimiter is not set (clean case) Added some logic to regenerate the Aboot cmdline when cmdline-aboot-end is set but duplicate parameters exists before (dirty case). Reorganized some code to handle duplicate parameter handling in the allowlist.	2022-07-04 11:01:03 -07:00
Stephen Sun	fe6be5da92	[202012] Configure different map between uplink and downlink on t1 switch in dual ToR scenario (#11299 ) - Why I did it Configure different DSCP_TO_TC_MAP between uplink and downlink on T1 switch in dual ToR scenario On T1 uplink, both DSCP 2/6 will be mapped to TC 1 for the purpose of avoiding such traffic occupying lossless buffers. On T1 downlink, they will be mapped to TC 2/6 respectively. (unchanged) - How I did it For vendors who want to configure different DSCP_TO_TC_MAP between uplinks and downlinks on T1, they should Define generate_dscp_to_tc_map macro in SKU's qos.json.j2 file Define map AZURE for downlink and AZURE_UPLINK for uplink Define jinja2 variable different_dscp_to_tc_map as True Signed-off-by: Stephen Sun <stephens@nvidia.com>	2022-07-03 15:58:06 +03:00
Stephen Sun	307d0e2aca	[Mellanox][202012] Support Mellanox-SN4600C-C64 as T1 switch in dual-ToR scenario (#11032 ) Why I did it Support Mellanox-SN4600C-C64 as T1 switch in dual-ToR scenario 1. Support additional queue and PG in buffer templates, including both traditional and dynamic model 2. Support mapping DSCP 2/6 to lossless traffic in the QoS template. 3. Add macros to generate additional lossless PG in the dynamic model 4. Adjust the order in which the generic/dedicated (with additional lossless queues) macros are checked and called to generate buffer tables in common template buffers_config.j2 - Buffer tables are rendered via using macros. - Both generic and dedicated macros are defined on our platform. Currently, the generic one is called as long as it is defined, which causes the generic one always being called on our platform. To avoid it, the dedicated macrio is checked and called first and then the generic ones. 5. Support MAP_PFC_PRIORITY_TO_PRIORITY_GROUP on ports with additional lossless queues. On Mellanox-SN4600C-C64, buffer configuration for t1 is calculated as: 40 * 100G downlink ports with 4 lossless PGs/queues, 1 lossy PG, and 3 lossy queues 16 * 100G uplink ports with 2 lossless PGs/queues, 1 lossy PG, and 5 lossy queues Signed-off-by: Stephen Sun stephens@nvidia.com How to verify it Run regression test.	2022-06-21 10:04:49 -07:00
bingwang-ms	6ddf5cd7dc	[202012] [cherry-pick] Generate switch level dscp_to_tc_map entry from qos_config template (#11132 ) * Generate switch level dscp_to_tc_map Signed-off-by: bingwang <wang.bing@microsoft.com>	2022-06-17 20:49:56 +08:00
Jing Kan	5b2261da37	Revert "[202012][openssh] openssh: Upgrade from 7.9 to 8.4, to match version in buster-backports (#10910 )" (#11136 ) This reverts commit `14fdcc815a`.	2022-06-17 20:46:43 +08:00
Saikrishna Arcot	044570c42e	Remove SSH host keys after installing the custom version of sshd (#10633 ) (#11140 ) * Remove SSH host keys after installing the custom version of sshd Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com> * Use an override for for sshd instead of overwriting the service file Don't overwrite upstream's .service file, and instead use an override file for making sure the host key(s) are generated. Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com>	2022-06-16 11:47:04 -07:00
Lukas Stockner	ab10005729	[swss] Clear VXLAN tunnel table from State DB on startup (#11078 ) *Clear VXLAN tunnel table from State DB on startup Signed-off-by: Lukas Stockner <lstockner@genesiscloud.com>	2022-06-10 11:50:56 -07:00
shlomibitton	2a9aa0836c	[202012] [Mellanox] [pmon] Fix for PMON service not starting when restarting SWSS service after fast/warm reboot (#10902 ) - Why I did it Recent change to delay PMON service in case of fast/warm reboot introduce an issue when restarting only SWSS service after fast/warm reboot for Nvidia platform. Since the timer is triggered only when the system boot, in a scenario when the system is after a fast/warm reboot and the user restart SWSS service, as part of syncd.sh script, PMON service will stop but the timer will not start again. - How I did it On syncd.sh script, in case of fast/warm indication, check if pmon.timer is running. If it is running it means we are at the first boot and continue normally. If it is not running, meaning the service was restarted, start the timer to keep the system behavior consistent. - How to verify it Run fast/warm reboot. service swss restart. Observe PMON service starting.	2022-06-08 09:46:54 +03:00
mssonicbld	855ae0491f	[ci/build]: Upgrade SONiC package versions (#11051 ) Upgrade SONiC Versions #11051	2022-06-08 08:42:04 +08:00
Richard.Yu	8f3edde302	[202012][BRCM SAI 4.3.5.3-5] Update saibcm for pcbb feature (#10998 ) Support Tunnel PFC/pcbb feature on Broadcom platform. How to verify it Tested build target, successful make target/docker-syncd-brcm.gz manual run those tests after installing sai binary within image 20201231.67 on 7050CX3 (TD3) T0 DUT, all passed fib/test_fib.py vxlan/test_vxlan_decap.py fdb/test_fdb.py decap/test_decap.py pfcwd/test_pfcwd_all_port_storm.py acl/null_route/test_null_route_helper.py acl/test_acl.py vlan/test_vlan.py platform_tests/test_reboot.py Signed-off-by: richardyu-ms <richard.yu@microsoft.com>	2022-06-06 09:54:00 -07:00
bingwang-ms	e159998657	[202012][cherry-pick] Add two extra lossless queues for bounced back traffic (#10715 ) * Add extra lossless queues Signed-off-by: bingwang <bingwang@microsoft.com>	2022-06-04 19:25:02 +08:00

1 2 3 4 5 ...

1021 Commits