sonic-buildimage

Author	SHA1	Message	Date
mssonicbld	d24995223d	[ci/build]: Upgrade SONiC package versions (#12987 )	2022-12-08 02:13:33 +08:00
mssonicbld	2cd0fcb9df	[ci/build]: Upgrade SONiC package versions (#12925 )	2022-12-04 19:57:40 +08:00
mssonicbld	f683a8737a	[ci/build]: Upgrade SONiC package versions (#12914 )	2022-12-02 18:46:44 +08:00
mssonicbld	3c88bee836	[ci/build]: Upgrade SONiC package versions (#12827 )	2022-11-27 18:59:21 +08:00
bingwang-ms	47d7e5d0d2	[202012] Apply separated DSCP_TO_TC_MAP and TC_TO_QUEUE_MAP on dualtor (#12792 ) * Apply separated DSCP_TO_TC_MAP and TC_TO_QUEUE_MAP on dualtor	2022-11-23 21:49:00 +08:00
mssonicbld	6a74213b25	[ci/build]: Upgrade SONiC package versions (#12796 )	2022-11-23 18:48:57 +08:00
Lorne Long	3402094fd0	[Build] Use apt-get to predictably support dependency ordered configuration of lazy packages (#12164 ) Why I did it The current lazy installer relies on a filename sort for both unpack and configuration steps. When systemd services are configured [started] by multiple packages the order is by filename not by the declared package dependencies. This can cause the start order of services to differ between first-boot and subsequent boots. Declared systemd service dependencies further exacerbate the issue (e.g. blocking the first-boot script). The current installer leaves packages un-configured if the package dependency order does not match the filename order. This also fixes a trivial bug in [Build]: Support to use symbol links for lazy installation targets to reduce the image size #10923 where externally downloaded dependencies are duplicated across lazy package device directories. How I did it Changed the staging and first-boot scripts to use apt-get: dpkg -i /host/image-$SONIC_VERSION/platform/$platform/.deb becomes apt-get -y install /host/image-$SONIC_VERSION/platform/$platform/.deb when dependencies are detected during image staging. How to verify it Apt-get critical rules Add a Depends= to the control information of a package. Grep the syslog for rc.local between images and observe the configuration order of packages change.	2022-11-23 10:41:28 +00:00
mssonicbld	b3d4305adc	[ci/build]: Upgrade SONiC package versions (#12756 )	2022-11-20 19:00:06 +08:00
Devesh Pathak	acca64a8bd	[202012] Clear /etc/resolv.conf before building image (#12686 ) Why I did it nameserver and domain entries from build system fsroot gets into sonic image. How I did it Clear /etc/resolv.conf before building image How to verify it Built image with it and verified with install that /etc/resolv.conf is empty	2022-11-15 10:09:09 +08:00
mssonicbld	db770a9353	[ci/build]: Upgrade SONiC package versions (#12690 )	2022-11-13 20:29:24 +08:00
Lawrence Lee	275adc6691	[arp_update]: Fix hardcoded vlan (#12566 ) Typo in prior PR #11919 hardcodes Vlan name. Change command to use the $vlan variable instead Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2022-11-11 18:01:15 +00:00
mssonicbld	5d5339dd13	[ci/build]: Upgrade SONiC package versions (#12676 )	2022-11-11 20:31:36 +08:00
bingwang-ms	4f4f4cba21	[202012] Add lossy scheduler for queue 7 (#12600 ) * Add lossy scheduler for queue 7	2022-11-10 10:25:03 +08:00
mssonicbld	856e536659	[ci/build]: Upgrade SONiC package versions (#12655 )	2022-11-09 23:36:30 +08:00
mssonicbld	bf39fca99e	[ci/build]: Upgrade SONiC package versions (#12604 )	2022-11-06 20:16:03 +08:00
mssonicbld	8f80dc3a1b	[ci/build]: Upgrade SONiC package versions (#12583 )	2022-11-02 21:21:29 +08:00
mssonicbld	0f721d8d98	[ci/build]: Upgrade SONiC package versions (#12550 )	2022-10-30 20:05:48 +08:00
mssonicbld	eee839fddf	[ci/build]: Upgrade SONiC package versions (#12547 )	2022-10-28 21:24:57 +08:00
mssonicbld	713747dc05	[ci/build]: Upgrade SONiC package versions (#12507 )	2022-10-26 19:58:39 +08:00
Devesh Pathak	d45fe19576	Fix to improve hostname handling (#12064 ) * Fix to improve hostname handling If config_db.json is missing hostname entry, hostname-config.sh ends up deleting existing entry too and hostname changes to default 'localhost' * default hostname to 'sonic` if missing in config file	2022-10-26 05:48:16 +00:00
zitingguo-ms	bafbfb5a26	Pickup fix and make up BRCM SAI version to 4.3.7.1-6 (#12486 ) Signed-off-by: zitingguo-ms <zitingguo@microsoft.com> Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-10-26 09:52:48 +08:00
mssonicbld	d2d25ac5f5	[ci/build]: Upgrade SONiC package versions (#12449 )	2022-10-21 21:48:16 +08:00
zitingguo-ms	08d1d60ccb	Pick up fixes and make up BRCM SAI version to 4.3.7.1-3 (#12439 ) Signed-off-by: zitingguo-ms <zitingguo@microsoft.com> Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-10-19 12:18:48 +08:00
mssonicbld	678edcb90f	[ci/build]: Upgrade SONiC package versions (#12417 )	2022-10-16 21:21:35 +08:00
mssonicbld	bb2d0986e2	[ci/build]: Upgrade SONiC package versions (#12407 )	2022-10-14 22:02:13 +08:00
xumia	2955a8dc72	[202012] Change submodule path from Azure to sonic-net (#12312 ) Why I did it Change the path of sonic submodules that point to "Azure" to point to "sonic-net" How I did it Replace "Azure" with "sonic-net" on all relevant paths of sonic submodules	2022-10-13 23:30:37 +08:00
mssonicbld	dd1db72c1f	[ci/build]: Upgrade SONiC package versions (#12368 )	2022-10-13 15:35:28 +08:00
mssonicbld	54b0b8b557	[ci/build]: Upgrade SONiC package versions (#12355 )	2022-10-11 21:45:13 +08:00
mssonicbld	30adf36ce3	[ci/build]: Upgrade SONiC package versions (#12302 )	2022-10-09 18:25:20 +08:00
gechiang	1e6d63a412	[202012][BRCMSAI] 4.3.7.1-2 to back out a change that broke 4.3.7.1-1 (#12298 ) This is basically the same as previous PR: (#12275) With backing out a change that was breaking the build. Copying the same info from that PR here.	2022-10-06 21:25:34 -07:00
gechiang	9c9d902ede	[202012]BRCM SAI 4.3.7.1-1 pick up fix CS00012263713 (mirrored packet with extra VLAN Tag) (#12275 ) Pick up fix for CS00012263713 (mirrored packet with extra VLAN Tag) BRCM SAI 4.3.7.1-1 Preliminary tests look fine. BGP neighbors were all up with proper routes programmed interfaces are all up Manually ran the following test cases on 7050CX3 (TD3) T0 DUT and all passed: fib/test_fib.py acl/test_acl.py arp/test_neighbor_mac_noptf.py fdb/test_fdb.py decap/test_decap.py pc/test_lag_2.py pc/test_po_cleanup.py pc/test_po_update.py everflow/test_everflow_ipv6.py everflow/test_everflow_testbed.py route/test_default_route.py ipfwd/test_dip_sip.py copp/test_copp.py crm/test_crm.py	2022-10-05 09:40:55 -07:00
mssonicbld	37a189b3f3	[ci/build]: Upgrade SONiC package versions (#12277 )	2022-10-05 20:57:41 +08:00
mssonicbld	7474a1bcd8	[ci/build]: Upgrade SONiC package versions (#12267 )	2022-10-04 21:20:48 +08:00
Ying Xie	64ce6696bb	[mux] skip mux operations during warm shutdown (#11937 ) * [mux] skip mux operations during warm shutdown - Enhance write_standby.py script to skip actions during warm shutdown. - Expand the support to BGP service. - MuX support was added by a previous PR. - don't skip action during warm recovery Signed-off-by: Ying Xie <ying.xie@microsoft.com>	2022-10-03 22:30:55 +00:00
Longxiang Lyu	893391f76e	[mux] Exit to write `standby` state to `active-active` ports (#11821 ) [mux] Exit to write standby state to `active-active` ports Signed-off-by: Longxiang Lyu <lolv@microsoft.com>	2022-10-03 19:52:51 +00:00
mssonicbld	a05d86a047	[ci/build]: Upgrade SONiC package versions (#12252 )	2022-10-02 18:38:46 +08:00
mssonicbld	f61c84b09e	[ci/build]: Upgrade SONiC package versions (#12243 )	2022-09-30 21:36:39 +08:00
mssonicbld	8994298788	[ci/build]: Upgrade SONiC package versions (#12199 )	2022-09-28 21:32:50 +08:00
mssonicbld	dc67ec2869	[ci/build]: Upgrade SONiC package versions (#12181 )	2022-09-25 20:10:04 +08:00
zitingguo-ms	95b19bbb46	Pick up fixes and make up BRCM SAI version to 4.3.7.1 (#12069 ) Signed-off-by: zitingguo-ms <zitingguo@microsoft.com> Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-09-22 22:59:51 -07:00
mssonicbld	2b5eb45d83	[ci/build]: Upgrade SONiC package versions (#12141 )	2022-09-21 20:37:27 +08:00
mssonicbld	4396d67d64	[ci/build]: Upgrade SONiC package versions (#12120 )	2022-09-20 19:51:45 +08:00
mssonicbld	6678e3356e	[ci/build]: Upgrade SONiC package versions (#12101 )	2022-09-18 19:48:13 +08:00
lixiaoyuner	0abf8d0419	Make client indentity by AME cert (#11946 ) * Make client indentity by AME cert * Join k8s cluster by ipv6 * Change join test cases * Test case bug fix * Improve read node label func * Configure kubelet and change test cases * For kubernetes version 1.22.2 * Fix undefine issue Signed-off-by: Yun Li <yunli1@microsoft.com>	2022-09-17 00:41:53 +00:00
mssonicbld	a195f8343f	[ci/build]: Upgrade SONiC package versions (#12068 )	2022-09-14 23:05:43 +08:00
mssonicbld	37ada92ad5	[ci/build]: Upgrade SONiC package versions (#12043 )	2022-09-11 18:59:58 +08:00
mssonicbld	5e8332f505	[ci/build]: Upgrade SONiC package versions (#12039 )	2022-09-09 21:46:28 +08:00
siqbal1986	10b6a5c402	[202012] StateDB table cleanup for VNET routes (#11999 ) * added cleanup of tables. 'VNET_ROUTE_TUNNEL_TABLE', 'VNET_ROUTE_TABLE'	2022-09-08 14:57:20 -07:00
mssonicbld	143af80061	[ci/build]: Upgrade SONiC package versions (#11976 )	2022-09-07 00:40:48 +08:00
mssonicbld	586a623422	[ci/build]: Upgrade SONiC package versions (#11911 )	2022-09-04 19:29:06 +08:00
Lawrence Lee	e821dd8551	[arp_update]: Set failed IPv6 neighbors to incomplete (#11919 ) After pinging any failed IPv6 neighbor entries, set the remaining failed/incomplete entries to a permanent INCOMPLETE state. This manual setting to INCOMPLETE prevents these entries from automatically transitioning to FAILED state, and since they are now incomplete any subsequent NA messages for these neighbors is able to resolve the entry in the cache. Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2022-09-02 21:57:47 +00:00
Ying Xie	4ab83170a5	[write_standby] update write_standby.py script (#11650 ) Why I did it The initial value has to be present for the state machines to work. In active-standby dual-tor scenario, or any hardware mux scenario, the value will be updtaed eventually with a delay. However, in active-active dual-tor scenario, there is no other mechanism to initialize the value and get state machines started. So this script will have to write something at start up time. For active-active dualtor, 'active' is a more preferred initial value, the state machine will switch the state to standby soon if link prober found link not in good state. How I did it Update the script to always provide initial values. How to verify it Tested on active-active dual-tor testbed. Signed-off-by: Ying Xie ying.xie@microsoft.com	2022-09-01 23:57:23 +00:00
Jing Zhang	9d3194c77a	Avoid write_standby in warm restart context (#11283 ) Avoid write_standby in warm restart context. sign-off: Jing Zhang zhangjing@microsoft.com Why I did it In warm restart context, we should avoid mux state change. How I did it Check warm restart flag before applying changes to app db. How to verify it Ran write_standby in table missing, key missing, field missing scenarios. Did a warm restart, app db changes were skipped. Saw this in syslog: WARNING write_standby: Taking no action due to ongoing warmrestart.	2022-09-01 23:57:17 +00:00
mssonicbld	ed68e4c97c	[ci/build]: Upgrade SONiC package versions (#11896 )	2022-08-30 22:44:47 +08:00
mssonicbld	347b2dddcd	[ci/build]: Upgrade SONiC package versions (#11757 )	2022-08-29 14:08:14 +08:00
mssonicbld	07082bb5f5	[ci/build]: Upgrade SONiC package versions (#11676 )	2022-08-16 13:07:32 +00:00
Stepan Blyshchak	8ab448a852	[swss.sh/syncd.sh] Trap only on EXIT (#11590 ) When using trap on SIGTERM the script will not react to the SIGTERM signal sent while a child is executing. I.e, the following script does not react on SIGTERM sent to it if it is waiting for sleep to finish: ``` trap "echo Handled SIGTERM" 0 2 3 15 echo "Before sleep" sleep inf echo "After sleep" ``` Instead, trap only on EXIT which covers also a scenario with exit on SIGINT, SIGTERM. Signed-off-by: Stepan Blyschak <stepanb@nvidia.com>	2022-08-11 20:38:20 +00:00
zitingguo-ms	5b5bd5e818	[202012 BRCM SAI 4.3.7.0] Pick up fixes and make up BRCM SAI version to 4.3.7.0 (#11681 ) Pick upfollowing fixes and update BRCM SAI to 4.3.7.0: CS00012208537: Add back previous commit 54c5bc4848eb748 CS00012253061,SONIC-63280: WB from 3.5 to 4.3, followed by WB to 4.3 CS00012207978: SDK-296517, time spent for SAI operations CS00012245601,SONIC-62898: Egress ACL Counted ad Interface TX drops Update pcbb with Fixes for CS00012243699 Upgrade on pcbb with Fixes for KB0025353, CS00012221689, CS00012221688, KB0025391, CS00012230519 commit of "CS00012221688:PFC frames egressing, PFC storm happens simultaneously on 2 ports" is purposely skipped to be picked up later due to SWSS dependency not ready. Why I did it How I did it How to verify it Tested build target, successful Manually run these tests after installing sai binary within image 20201231.73 on 7050CX3 (TD3) T0 DUT, all passed. vxlan/test_vxlan_decap.py fdb/test_fdb.py pfcwd/test_pfcwd_all_port_storm.py acl/null_route/test_null_route_helper.py acl/test_acl.py vlan/test_vlan.py platform_tests/test_reboot.py Signed-off-by: zitingguo-ms <zitingguo@microsoft.com>	2022-08-10 15:02:47 -07:00
Jing Zhang	ffd9e190e1	Update WARM START FINALIZER to wait for linkmgrd to reconcile (#11477 ) Spanning from sonic-net/sonic-linkmgrd#76, this PR is to update warm restart finalizer to wait for linkmgrd to be reconciled. sign-off: Jing Zhang zhangjing@microsoft.com Why I did it To make sure finalizer save config after linkmgrd's reconciliation. How I did it Add linkmgrd to the reconciliation wait list of warmboot finalizer. How to verify it Verified on lab device, linkmgrd reconciled as expected.	2022-08-09 21:05:12 +00:00
mssonicbld	14f93e15c6	[ci/build]: Upgrade SONiC package versions (#11629 ) Why I did it Upgrade SONiC Versions	2022-08-07 11:27:16 +08:00
Lawrence Lee	04ba6da1ab	[202012][arp_update]: Resolve failed neighbors on dualtor (#11641 ) In arp_update, check for FAILED or INCOMPLETE kernel neighbor entries and manually ping them to try and resolve the neighbor Signed-off-by: Lawrence Lee <lawlee@microsoft.com>	2022-08-05 23:30:04 -07:00
tjchadaga	6d66d9b8fc	Revert "Add load_minigraph option to include traffic-shift-away during config migration (#11403 )" (#11625 ) This reverts commit `6c2f99a327`.	2022-08-06 10:05:45 +05:30
bingwang-ms	84aca00847	[202012]Support different `DSCP_TO_TC_MAP` for T1 in dualtor deployment (#11580 ) Why I did it This PR is to backport #11569 into 202012 branch. This PR is to apply different DSCP_TO_TC_MAP to downlink and uplink ports on T1 in dualtor deployment. For T1 downlink ports (To T0) The DSCP_TO_TC_MAP is not changed. DSCP2 and DSCP6 are mapped to TC2 and TC6 respectively. For T1 uplink ports (To T1) A new DSCP_TO_TC_MAP\|AZURE_UPLINK is defined and applied. DSCP2 and DSCP6 are mapped to TC1 to avoid mixing up lossy and lossless traffic from T2. The extra lossy PG2 and PG6 added in PR #11157 is reverted as well because no traffic from T2 is mapped to PG2 or PG6 now. How I did it Define a new map DSCP_TO_TC_MAP\|AZURE_UPLINK for 7260 T1. How to verify it Verified by test case in test_j2files.py.	2022-08-01 08:59:45 -07:00
Nikola Dancejic	c5a5734242	[swss] Adding bgp container as dependent of swss (#11168 ) What I did: Added bgp as a dependent of swss Why I did it: bgp container was not restarting on swss crash. When swss crashes, linkmgrd doesn't initate a switchover because it cannot access the default route from orchagent. Bringing down bgp with swss will isolate the ToR, causing linkmgrd to initiate a switchover to the peer ToR avoiding significant packet loss. Signed-off-by: Nikola Dancejic <ndancejic@microsoft.com>	2022-07-29 09:37:09 -07:00
Lior Avramov	a40aca43b9	[memory_checker] Do not check memory usage of containers if docker daemon is not running (#11476 ) Fix in Monit memory_checker plugin. Skip fetching running containers if docker engine is down (can happen in deinit). This PR fixes issue #11472. Signed-off-by: liora liora@nvidia.com Why I did it In the case where Monit runs during deinit flow, memory_checker plugin is fetching the running containers without checking if Docker service is still running. I added this check. How I did it Use systemctl is-active to check if Docker engine is still running. How to verify it Use systemctl to stop docker engine and reload Monit, no errors in log and relevant print appears in log. Which release branch to backport (provide reason below if selected) The fix is required in 202205 and 202012 since the PR that introduced the issue was cherry picked to those branches (#11129).	2022-07-27 23:28:19 +00:00
tjchadaga	6c2f99a327	Add load_minigraph option to include traffic-shift-away during config migration (#11403 )	2022-07-27 23:27:21 +00:00
bingwang-ms	c5eb031111	[202012] Add flag to control the generation of global level map (#11451 ) Why I did it This PR is to cherry-pick #11448 to 202012 branch after resolving conflicts. There are conflicts in files/build_templates/qos_config.j2 src/sonic-config-engine/tests/test_j2files.py	2022-07-15 09:44:45 -07:00
mssonicbld	550ab26fc7	[ci/build]: Upgrade SONiC package versions (#11422 )	2022-07-12 15:39:32 +00:00
Neetha John	26ee4ae4a4	Add backend acl template (#11220 ) Why I did it Storage backend has all vlan members tagged. If untagged packets are received on those links, they are accounted as RX_DROPS which can lead to false alarms in monitoring tools. Using this acl to hide these drops. How I did it Created a acl template which will be loaded during minigraph load for backend. This template will allow tagged vlan packets and dropped untagged How to verify it Unit tests Signed-off-by: Neetha John <nejo@microsoft.com>	2022-07-08 21:39:39 +00:00
mssonicbld	9a86fa9264	[ci/build]: Upgrade SONiC package versions (#11074 ) Upgrade SONiC Versions	2022-07-06 11:00:50 +08:00
xumia	32cda89f93	[Build]: Support to use symbol links for lazy installation targets to reduce the image size (#10923 ) Why I did it Support to use symbol links in platform folder to reduce the image size. The current solution is to copy each lazy installation targets (xxx.deb files) to each of the folders in the platform folder. The size will keep growing when more and more packages added in the platform folder. For cisco-8000 as an example, the size will be up to 2G, while most of them are duplicate packages in the platform folder. How I did it Create a new folder in platform/common, all the deb packages are copied to the folder, any other folders where use the packages are the symbol links to the common folder. Why platform.tar? We have implemented a patch for it, see #10775, but the problem is the the onie use really old unzip version, cannot support the symbol links. The current solution is similar to the PR 10775, but make the platform folder into a tar package, which can be supported by onie. During the installation, the package.tar will be extracted to the original folder and removed.	2022-07-05 20:57:49 +00:00
yozhao101	4487a962e3	[memory_checker] Do not check memory usage of containers which are not created (#11129 ) Signed-off-by: Yong Zhao yozhao@microsoft.com Why I did it This PR aims to fix an issue (#10088) by enhancing the script memory_checker. Specifically, if container is not created successfully during device is booted/rebooted, then memory_checker do not need check its memory usage. How I did it In the script memory_checker, a function is added to get names of running containers. If the specified container name is not in current running container list, then this script will exit without checking its memory usage. How to verify it I tested on a lab device by following the steps: Stops telemetry container with command sudo systemctl stop telemetry.service Removes telemetry container with command docker rm telemetry Checks whether the script memory_checker ran by Monit will generate the syslog message saying it will exit without checking memory usage of telemetry.	2022-07-05 20:57:45 +00:00
Samuel Angebault	d15a484dfa	[202012][Arista] Fix cmdline generation during warm-reboot from 201811/201911 (#11161 ) Issue fixed: when performing a warm-reboot or fast-reboot from 201811 or 201911 to 202012 the kernel command line contains duplicate information. This issue is related to a change that was made to make 202012 boot0 file more futureproof. A cold reboot brings everything back into a clean slate though not always desirable. Changes done: Added some logic to properly detect the end of the Aboot cmdline when cmdline-aboot-end delimiter is not set (clean case) Added some logic to regenerate the Aboot cmdline when cmdline-aboot-end is set but duplicate parameters exists before (dirty case). Reorganized some code to handle duplicate parameter handling in the allowlist.	2022-07-04 11:01:03 -07:00
Stephen Sun	fe6be5da92	[202012] Configure different map between uplink and downlink on t1 switch in dual ToR scenario (#11299 ) - Why I did it Configure different DSCP_TO_TC_MAP between uplink and downlink on T1 switch in dual ToR scenario On T1 uplink, both DSCP 2/6 will be mapped to TC 1 for the purpose of avoiding such traffic occupying lossless buffers. On T1 downlink, they will be mapped to TC 2/6 respectively. (unchanged) - How I did it For vendors who want to configure different DSCP_TO_TC_MAP between uplinks and downlinks on T1, they should Define generate_dscp_to_tc_map macro in SKU's qos.json.j2 file Define map AZURE for downlink and AZURE_UPLINK for uplink Define jinja2 variable different_dscp_to_tc_map as True Signed-off-by: Stephen Sun <stephens@nvidia.com>	2022-07-03 15:58:06 +03:00
Stephen Sun	307d0e2aca	[Mellanox][202012] Support Mellanox-SN4600C-C64 as T1 switch in dual-ToR scenario (#11032 ) Why I did it Support Mellanox-SN4600C-C64 as T1 switch in dual-ToR scenario 1. Support additional queue and PG in buffer templates, including both traditional and dynamic model 2. Support mapping DSCP 2/6 to lossless traffic in the QoS template. 3. Add macros to generate additional lossless PG in the dynamic model 4. Adjust the order in which the generic/dedicated (with additional lossless queues) macros are checked and called to generate buffer tables in common template buffers_config.j2 - Buffer tables are rendered via using macros. - Both generic and dedicated macros are defined on our platform. Currently, the generic one is called as long as it is defined, which causes the generic one always being called on our platform. To avoid it, the dedicated macrio is checked and called first and then the generic ones. 5. Support MAP_PFC_PRIORITY_TO_PRIORITY_GROUP on ports with additional lossless queues. On Mellanox-SN4600C-C64, buffer configuration for t1 is calculated as: 40 * 100G downlink ports with 4 lossless PGs/queues, 1 lossy PG, and 3 lossy queues 16 * 100G uplink ports with 2 lossless PGs/queues, 1 lossy PG, and 5 lossy queues Signed-off-by: Stephen Sun stephens@nvidia.com How to verify it Run regression test.	2022-06-21 10:04:49 -07:00
bingwang-ms	6ddf5cd7dc	[202012] [cherry-pick] Generate switch level dscp_to_tc_map entry from qos_config template (#11132 ) * Generate switch level dscp_to_tc_map Signed-off-by: bingwang <wang.bing@microsoft.com>	2022-06-17 20:49:56 +08:00
Jing Kan	5b2261da37	Revert "[202012][openssh] openssh: Upgrade from 7.9 to 8.4, to match version in buster-backports (#10910 )" (#11136 ) This reverts commit `14fdcc815a`.	2022-06-17 20:46:43 +08:00
Saikrishna Arcot	044570c42e	Remove SSH host keys after installing the custom version of sshd (#10633 ) (#11140 ) * Remove SSH host keys after installing the custom version of sshd Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com> * Use an override for for sshd instead of overwriting the service file Don't overwrite upstream's .service file, and instead use an override file for making sure the host key(s) are generated. Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com>	2022-06-16 11:47:04 -07:00
Lukas Stockner	ab10005729	[swss] Clear VXLAN tunnel table from State DB on startup (#11078 ) *Clear VXLAN tunnel table from State DB on startup Signed-off-by: Lukas Stockner <lstockner@genesiscloud.com>	2022-06-10 11:50:56 -07:00
shlomibitton	2a9aa0836c	[202012] [Mellanox] [pmon] Fix for PMON service not starting when restarting SWSS service after fast/warm reboot (#10902 ) - Why I did it Recent change to delay PMON service in case of fast/warm reboot introduce an issue when restarting only SWSS service after fast/warm reboot for Nvidia platform. Since the timer is triggered only when the system boot, in a scenario when the system is after a fast/warm reboot and the user restart SWSS service, as part of syncd.sh script, PMON service will stop but the timer will not start again. - How I did it On syncd.sh script, in case of fast/warm indication, check if pmon.timer is running. If it is running it means we are at the first boot and continue normally. If it is not running, meaning the service was restarted, start the timer to keep the system behavior consistent. - How to verify it Run fast/warm reboot. service swss restart. Observe PMON service starting.	2022-06-08 09:46:54 +03:00
mssonicbld	855ae0491f	[ci/build]: Upgrade SONiC package versions (#11051 ) Upgrade SONiC Versions #11051	2022-06-08 08:42:04 +08:00
Richard.Yu	8f3edde302	[202012][BRCM SAI 4.3.5.3-5] Update saibcm for pcbb feature (#10998 ) Support Tunnel PFC/pcbb feature on Broadcom platform. How to verify it Tested build target, successful make target/docker-syncd-brcm.gz manual run those tests after installing sai binary within image 20201231.67 on 7050CX3 (TD3) T0 DUT, all passed fib/test_fib.py vxlan/test_vxlan_decap.py fdb/test_fdb.py decap/test_decap.py pfcwd/test_pfcwd_all_port_storm.py acl/null_route/test_null_route_helper.py acl/test_acl.py vlan/test_vlan.py platform_tests/test_reboot.py Signed-off-by: richardyu-ms <richard.yu@microsoft.com>	2022-06-06 09:54:00 -07:00
bingwang-ms	e159998657	[202012][cherry-pick] Add two extra lossless queues for bounced back traffic (#10715 ) * Add extra lossless queues Signed-off-by: bingwang <bingwang@microsoft.com>	2022-06-04 19:25:02 +08:00
bingwang-ms	7ec6a60230	[cherry-pick] [202012] Update qos config to clear queues for bounced back traffic (#10608 ) * Update qos config to clear queues for bounced back traffic Signed-off-by: bingwang <wang.bing@microsoft.com>	2022-06-02 16:29:25 +08:00
Jing Kan	14fdcc815a	[202012][openssh] openssh: Upgrade from 7.9 to 8.4, to match version in buster-backports (#10910 ) * Use buster-backports version * Use dget dsc file instead source repo * Update make files * Upgrade openssh-client to 8.4 in base image * Remove useless installation * Install openssh-server from buster-backports in build_debian * Update dev buster package version list Signed-off-by: Jing Kan jika@microsoft.com	2022-06-02 16:06:22 +08:00
xumia	06addae853	Revert "Reduce image size for lazy installation packages (#10775 )" (#10916 ) This reverts commit `15cf9b0d70`. Why I did it Revert the PR #10775, for it has impact on onie installation. It is caused by the symbol links not supported in some of the onie unzip. We will enable after fixing the issue, see #10914	2022-05-27 17:00:50 +00:00
mssonicbld	8ce3cab508	[ci/build]: Upgrade SONiC package versions (#10732 ) Co-authored-by: mssonicbld <vsts@fv-az31-361.b1uo4dmaffwenkazr3a2h2ovdb.jx.internal.cloudapp.net>	2022-05-27 01:10:32 -07:00
shlomibitton	c71c91e2b0	[202012] [Fastboot] Delay PMON service for better fastboot performance (#10745 ) #### Why I did it Profiling the system state on init after fast-reboot during create_switch function execution, it is possible to see few python scripts running at the same time. This parallel execution consume CPU time and the duration of create_switch is longer than it should be. Following this finding, and the motivation to ensure these services will not interfere in the future, PMON is delayed in 90 seconds until the system finish the init flow after fastboot. #### How I did it Add a timer for PMON service. Exclude for MLNX platform the start trigger of PMON when SYNCD starts in case of fastboot. Copy the timer file to the host bin image. #### How to verify it Run fast-reboot on MLNX platform and observe faster create_switch execution time.	2022-05-15 23:31:32 -07:00
shlomibitton	bca8a244c6	[202012] [Fastboot] Delay LLDP service for better fastboot performance (#10568 ) (#10744 ) This PR is to backport a fix #10568 This PR is dependent on PR: #10745 - Why I did it Profiling the system state on init after fast-reboot during create_switch function execution, it is possible to see few python scripts running at the same time. This parallel execution consume CPU time and the duration of create_switch is longer than it should be. Following this finding, and the motivation to ensure these services will not interfere in the future, LLDP is delayed in 90 seconds until the system finish the init flow after fastboot. - How I did it Add a timer for LLDP service. Copy the timer file to the host bin image. - How to verify it Run fast-reboot on MLNX platform and observe faster create_switch execution time.	2022-05-15 15:05:29 +03:00
Junchao-Mellanox	4f326e8779	Fix race condition between networking service and interface-config service (#10573 ) (#10766 ) Backport https://github.com/Azure/sonic-buildimage/pull/10573 to 202012. #### Why I did it The PR is aimed to fix a bug that mgmt port eth0 may loss IP even if user configured static IP of eth0. This is not a always reproduceable issue, the reproducing flow is like: 1. Systemd starts networking service, which runs a dhcp based configuration and assigned an ip from dhcp. 2. Systemd starts interface-config service who depends on networking service 3. Interface-config service runs command “ifdown –force eth0”, check [line](`16717d2dc5/files/image_config/interfaces/interfaces-config.sh (L4)`). but networking service is still running so that this [line](`ac32bec0e2/ifupdown2/ifupdown/main.py (L74)`) failed with error: “error: Another instance of this program is already running.”. This error is printed by ifupdown2 lib who is the main process of networking service. So, ifdown actually does not work here, the ip of eth0 is not down. 4. Interface-config service updates /etc/networking/interface to static configuration. 5. Interface-config service runs command “systemctl restart networking”. This command kills the previous networking related processes (log: networking.service: Main process exited, code=killed, status=15/TERM), and try to reconfigure the ip address with static configuration. But it detects that the configured IP and the existing IP are the same, and it does not really configure the ip to kernel. Hence, the ip is still getting from dhcp. (this could be a bug of ifupdown2: previous ip is from dhcp, new ip is a static ip, it treats them as same instead of re-configuring the IP) 6. When the lease of the ip expires, the ip of eth0 is removed by kernel and the issue reproduces. The issue is not always reproduceable because networking service usually runs fast so that it won't hit step#3. #### How I did it Check networking service state before running "ifdown –force eth0", wait for it done if it is activating. #### How to verify it Manual test.	2022-05-14 14:58:24 -07:00
xumia	951d93e362	Reduce image size for lazy installation packages (#10775 ) Why I did it The image size is too large, when there are multiple lazy packages and multiple platforms. It is not necessary to keep the lazy installation packages in multiple copies. For cisco image, the image size will reduce from 3.5G to 1.7G. How I did it Use symbol links to only keep one package for each of the lazy package. Make a new folder fsroot/platform/common Copy the lazy packages into the folder. When using a package in each of the platform, such as x86_64-grub, x86_64-8800_rp-r0, x86_64-8201_on-r0, etc, only make a symbol link to the package in the common folder.	2022-05-10 06:44:40 +00:00
Samuel Angebault	705d3c0804	[Arista] Remove arista.log from rsyslog default logrotate (#9731 ) Why I did it In parallel of this change Arista added a custom logrotate configuration as part of its driver library. Having 2 logrotate configuration for the same log file triggers an issue. Fixes aristanetworks/sonic#38 How I did it Arista merged a few changes in sonic-buildimage which added a logrotate configuration aristanetworks/sonic@e43c797 It is therefore the right path to remove the arista.log line from the logrotate.d/rsyslog configuration. How to verify it Logrotate works without any error message, arista log rotation happens and arista daemons still append logs once file was truncated.	2022-04-28 23:58:41 +00:00
mssonicbld	1c9cdc4c7a	[ci/build]: Upgrade SONiC package versions (#10594 )	2022-04-27 15:25:14 +00:00
yozhao101	e6c18fa6dd	[Monit] Fix the issue which shows Monit can not reset its counter. (#10288 ) Signed-off-by: Yong Zhao <yozhao@microsoft.com> Why I did it This PR aims to fix the Monit issue which shows Monit can't reset its counter when monitoring memory usage of telemetry container. Specifically the Monit configuration file related to monitoring memory usage of telemetry container is as following: check program container_memory_telemetry with path "/usr/bin/memory_checker telemetry 419430400" if status == 3 for 10 times within 20 cycles then exec "/usr/bin/restart_service telemetry" If memory usage of telemetry container is larger than 400MB for 10 times within 20 cycles (minutes), then it will be restarted. Recently we observed, after telemetry container was restarted, its memory usage continuously increased from 400MB to 11GB within 1 hour, but it was not restarted anymore during this 1 hour sliding window. The reason is Monit can't reset its counter to count again and Monit can reset its counter if and only if the status of monitored service was changed from Status failed to Status ok. However, during this 1 hour sliding window, the status of monitored service was not changed from Status failed to Status ok. Currently for each service monitored by Monit, there will be an entry showing the monitoring status, monitoring mode etc. For example, the following output from command sudo monit status shows the status of monitored service to monitor memory usage of telemetry: Program 'container_memory_telemetry' status Status ok monitoring status Monitored monitoring mode active on reboot start last exit value 0 last output - data collected Sat, 19 Mar 2022 19:56:26 Every 1 minute, Monit will run the script to check the memory usage of telemetry and update the counter if memory usage is larger than 400MB. If Monit checked the counter and found memory usage of telemetry is larger than 400MB for 10 times within 20 minutes, then telemetry container was restarted. Following is an example status of monitored service: Program 'container_memory_telemetry' status Status failed monitoring status Monitored monitoring mode active on reboot start last exit value 0 last output - data collected Tue, 01 Feb 2022 22:52:55 After telemetry container was restarted. we found memory usage of telemetry increased rapidly from around 100MB to more than 400MB during 1 minute and status of monitored service did not have a chance to be changed from Status failed to Status ok. How I did it In order to provide a workaround for this issue, Monit recently introduced another syntax format repeat every <n> cycles related to exec. This new syntax format will enable Monit repeat executing the background script if the error persists for a given number of cycles. How to verify it I verified this change on lab device str-s6000-acs-12. Another pytest PR (Azure/sonic-mgmt#5492) is submitted in sonic-mgmt repo for review.	2022-04-21 22:00:42 +00:00
Samuel Angebault	9de6b2ca12	[Arista] Fix arista-net initramfs hook (#10626 ) The interface renaming logic fails if one interface is missing. Because of the `set -e` the whole initramfs hook would abort early on error. This change fixes the current behavior to make sure missing interfaces are properly skipped and ensure existing interface are renamed.	2022-04-20 10:03:37 -07:00
Jing Kan	4ee75f490e	[202012][copp_cfg] Enable dhcp trap for BmcMgmtToRRouter (#10596 ) Signed-off-by: Jing Kan jika@microsoft.com	2022-04-19 15:59:20 +08:00
Stepan Blyshchak	fa1e364f54	[services] kill container on stop in warm/fast mode (#10511 ) To optimize stop on warm boot, added kill for containers Use service "kill" in the shutdown path for fast and warm reboot. For all other reload methods, service "stop" is used. This is done to save time in shutdown path, and to overall improve the time spent in warm and fast reload. How - Use service_mgmt.sh to trigger common logic to initiate kill (fast/warm) or stop (cold) for database.sh, radv.sh, snmp.sh, telemetry.sh, mgmt-framework.sh Signed-off-by: Stepan Blyschak <stepanb@nvidia.com>, Vaibhav H D <vaibhav.dixit@microsoft.com>	2022-04-18 14:27:48 -07:00
Ying Xie	6af3de4372	[202012][copp cfg] enable dhcp trap for a couple more devices (#10582 ) * [copp cfg] enable copp trap for a couple more devices Signed-off-by: Ying Xie <ying.xie@microsoft.com>	2022-04-15 11:47:02 -07:00
Saikrishna Arcot	29b6f62902	[202012] Run tune2fs during initramfs instead of image install (#10558 ) If it is run during image install, it's not guaranteed that the installation environment will have tune2fs available. Therefore, run it during initramfs instead. Signed-off-by: Saikrishna Arcot <sarcot@microsoft.com>	2022-04-12 19:59:24 -07:00
mssonicbld	e0fa07307a	[ci/build]: Upgrade SONiC package versions (#10395 ) [ci/build]: Upgrade SONiC package versions (#10395)	2022-04-10 17:00:00 +08:00

1 2 3 4 5 ...

1104 Commits