linkerd2

Commit Graph

Author	SHA1	Message	Date
Alex Leong	755538b84a	Resolve gateway hostnames into IP addresses (#4588 ) Fixes #4582 When a target cluster gateway is exposed as a hostname rather than with a fixed IP address, the service mirror controller fails to create mirror services and gateway mirrors for that gateway. This is because we only look at the IP field of the gateway service. We make two changes to address this problem: First, when extracting the gateway spec from a gateway that has a hostname instead of an IP address, we do a DNS lookup to resolve that hostname into an IP address to use in the mirror service endpoints and gateway mirror endpoints. Second, we schedule a repair job on a regular (1 minute) to update these endpoint objects. This has the effect of re-resolving the DNS names every minute to pick up any changes in DNS resolution. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-06-15 10:33:49 -07:00
Zahari Dichev	f01bcfe722	Tweak service-mirror log levels (#4562 ) This PR just modifies the log levels on the probe and cluster watchers to emit in INFO what they would emit in DEBUG. I think it makes sense as we need that information to track problems. The only difference is that when probing gateways we only log if the probe attempt was unsuccessful. Fix #4546	2020-06-05 13:12:36 -07:00
Zahari Dichev	3365455e45	Fix mc labels (#4560 ) Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-06-05 19:36:09 +03:00
Zahari Dichev	b6b95455aa	Fix load balancer missing ip race condition (#4554 ) Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-06-05 19:35:47 +03:00
Alex Leong	cffa07ddba	Update gateway identity on gateway mirror endpoints (#4559 ) When the identity annotation on a gateway service is updated, this change is not propagated to the mirror gateway endpoints object. This is because the annotations are updated on the wrong object and the changes are lost. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-06-05 09:21:35 -07:00
Alex Leong	0f84ff61db	Update gateway mirror ports (#4551 ) * Update gateway mirror spec when remote gateway changes Signed-off-by: Alex Leong <alex@buoyant.io> * Only update ports Signed-off-by: Alex Leong <alex@buoyant.io>	2020-06-04 17:25:46 +03:00
Kevin Leimkuhler	8a932ac905	Change text to use source/target terminology in events and metrics (#4527 ) Change terminology from local/remote to source/target in events and metrics. This does not change any variable, function, struct, or field names since testing is still improving Signed-off-by: Kevin Leimkuhler <kevin@kleimkuhler.com>	2020-06-03 15:02:39 -04:00
Oliver Gould	7cc5e5c646	multicluster: Use the proxy as an HTTP gateway (#4528 ) This change modifies the linkerd-gateway component to use the inbound proxy, rather than nginx, for gateway. This allows us to detect loops and propagate identity through the gateway. This change also cleans up port naming to `mc-gateway` and `mc-probe` to resolve conflicts with Kubernetes validation. --- * proxy: v2.99.0 The proxy can now operate as gateway, routing requests from its inbound proxy to the outbound proxy, without passing the requests to a local application. This supports Linkerd's multicluster feature by adding a `Forwarded` header to propagate the original client identity and assist in loop detection. --- * Add loop detection to inbound & TCP forwarding (linkerd/linkerd2-proxy#527) * Test loop detection (linkerd/linkerd2-proxy#532) * fallback: Unwrap errors recursively (linkerd/linkerd2-proxy#534) * app: Split inbound/outbound constructors into components (linkerd/linkerd2-proxy#533) * Introduce a gateway between inbound and outbound (linkerd/linkerd2-proxy#540) * gateway: Add a Forwarded header (linkerd/linkerd2-proxy#544) * gateway: Return errors instead of responses (linkerd/linkerd2-proxy#547) * Fail requests that loop through the gateway (linkerd/linkerd2-proxy#545) * inject: Support config.linkerd.io/enable-gateway This change introduces a new annotation, config.linkerd.io/enable-gateway, that, when set, enables the proxy to act as a gateway, routing all traffic targetting the inbound listener through the outbound proxy. This also removes the nginx default listener and gateway port of 4180, instead using 4143 (the inbound port). * proxy: v2.100.0 This change modifies the inbound gateway caching so that requests may be routed to multiple leaves of a traffic split. --- * inbound: Do not cache gateway services (linkerd/linkerd2-proxy#549)	2020-06-02 19:37:14 -07:00
Kevin Leimkuhler	d7f84e6c7b	Change help text to use source/target terminology in service-mirror and healthchecks (#4524 ) Change terminology from local/remote to source/target in service-mirror and healthchecks help text. This does not change any variable, function, struct, or field names since testing is still improving Signed-off-by: Kevin Leimkuhler <kevin@kleimkuhler.com>	2020-06-02 15:21:52 -04:00
Alex Leong	91a067c924	Rename gateway ports (#4526 ) * Rename gateway ports Signed-off-by: Alex Leong <alex@buoyant.io> * fmt Signed-off-by: Alex Leong <alex@buoyant.io>	2020-06-02 09:08:23 +03:00
Kevin Leimkuhler	b4804a0bb5	Format fix (#4525 ) Fixes CI failures Signed-off-by: Kevin Leimkuhler <kevin@kleimkuhler.com>	2020-06-01 18:51:00 -04:00
Zahari Dichev	6c3922a7f1	Probe manager simplification (#4510 ) There are a few notable things happening in this PR: - the probe manager has been decoupled from the cluster_watcher. Now its only responsibility is to watch for mirrored gateways beeing created and to probe them. This means that probes are initiated for all gateways no matter whether there are mirrored services being paired - the number of paired services is derived from the existing services in the cluster rather than being published as a metric by the prober - there are no events being exchanged between the cluster watcher and the probe manager Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-06-01 14:41:29 -07:00
Zahari Dichev	f7f70690fb	Fix resync bug + service selection annotations (#4453 ) THis PR addresses two problems: - when a resync happens (or the mirror controller is restarted) we incorrectly classify the remote gateway as a mirrored service that is not mirrored anymore and we delete it - when updating services due to a gateway update, we need to select only the services for the particular cluster The latter fixes #4451	2020-05-21 14:15:13 -07:00
Alex Leong	acacf2e023	Add --close-wait-timeout inject flag (#4409 ) Depends on https://github.com/linkerd/linkerd2-proxy-init/pull/10 Fixes #4276 We add a `--close-wait-timeout` inject flag which configures the proxy-init container to run with `privileged: true` and to set `nf_conntrack_tcp_timeout_close_wait`. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-05-21 14:14:14 -07:00
Alex Leong	9cd4557644	Properly show the meshed count for non-selector services (#4446 ) When viewing the output of `linkerd stat` for services which do not have a selector (such as services created by the service-mirror, for example) the meshed count column shows the total number which exist, even though the service actually selects no pods at all. We update the StatSummary implementation to account for services which have no selector. Additionally, we update the logic of the `--unmeshed` flag. When the `--unmeshed` flag is not set, we typically skip rows for unmeshed resources because those resources would have no stats. This is not appropriate to do when the `--from` flag is also set because in this case, metrics are not collected on the target resource but are instead collected on the client-side. This means that stats can be present, even for unmeshed resources and these resources should still be displayed, even if the `--unmeshed` flag is not set. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-05-20 10:08:27 -07:00
Zahari Dichev	31e33d18d3	Enable service mirroring to work in private networks (#4440 ) This change creates a gateway proxy for every gateway. This enables the probe worker to leverage the destination service functionality in order to discover the identity of the gateway. Fix #4411 Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-20 19:48:36 +03:00
Zahari Dichev	6574f124a7	Restrict Service mirror RBACs (#4426 ) This PR introduces a few changes that were requested after a bit of service mirror reviewing. - we restrict the RBACs so the service mirror controller cannot read secrets in all namespaces but only in the one that it is installed in - we unify the namespace namings so all multicluster resources are installedi n `linkerd-multicluster` on both clusters - fixed checks to account for changes Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-20 17:08:01 +03:00
Zahari Dichev	4176580a0f	Threadsafe buffering listener (#4359 ) * Add thread safety to watcher tests Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-14 20:45:41 +03:00
Zahari Dichev	115bab9868	Fix gateway update problems (#4388 ) * Fix gateway update problems Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-14 10:59:30 -05:00
Zahari Dichev	ef1a2c2b10	Multicluster dashboard for traffic metrics (#4178 ) This change adds labels to endpoints that target remote services. It also adds a Grafana dashboard that can be used to monitor multicluster traffic. Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-14 17:48:27 +03:00
Zahari Dichev	fd59ce532d	Add better logging to service mirror controller (#4361 ) Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-11 10:30:16 +03:00
Zahari Dichev	edd9b654a7	Make gateway require TLS for incoming requests (#4339 ) Make gateway require TLS for incoming requests Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-11 10:07:48 +03:00
Alex Leong	8fbaa3ef9b	Don't send NoEndpoints during pod updates for ip watches (#4338 ) When the proxy has an IP watch on a pod and the destination controller gets a pod update event, the destination controller sends a NoEndpoints message to all listeners followed by an Add with the new pod state. This can result in the proxy's load balancer being briefly empty and could result in failing requests in the period. Since consecutive Add events with the same address will override each other, we can simply send the Adds without needing to clear the previous state with a NoEndpoints message.	2020-05-07 16:10:17 -07:00
Zahari Dichev	4e82ba8878	Multicluster checks (#4279 ) Multicluster checks Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-05 10:19:38 +03:00
Zahari Dichev	cd04b94bb9	Probe manager events emission tests (#4312 ) Probe manager events emission tests Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-05-05 08:57:05 +03:00
Alex Leong	40b921508f	Inject LINKERD2_PROXY_DESTINATION_GET_NETWORKS proxy variable (#4300 ) Fixes #3807 By setting the LINKERD2_PROXY_DESTINATION_GET_NETWORKS environment variable, we configure the Linkerd proxy to do destination lookups for authorities which are IP addresses in the private network range. This allows us to get destination metadata including identity for HTTP requests which target an IP address in the cluster, Prometheus metrics scrape requests, for example. This change allowed us to update the "direct edges" test which ensures that the edges command produces correct output for traffic which is addressed directly to a pod IP. We also re-enabled the "linkerd stat" integration tests which had been disabled while the destination service did not yet support these types of IP queries. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-04-30 11:22:24 -07:00
Zahari Dichev	5149152ef3	Multicluster gateway and remote setup command (#4265 ) Add multicluster gateway and setup command Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-04-29 20:33:23 +03:00
Zahari Dichev	17dacf5548	Add gateways command, allowing the retrieval of gateway stats (#4241 ) Add gateways command, allowing the retrieval of gateway stats Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-04-27 13:55:01 +03:00
Zahari Dichev	09262ebd72	Add liveliness checks and metrics for multicluster gateway (#4233 ) Add liveliness checks for gateway Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-04-27 13:06:58 +03:00
Tarun Pothulapati	2b1cbc6fc1	charts: Using downwardAPI to mount labels to the proxy container (#4199 ) * use downward API to mount labels to the proxy container as a volume * add namespace as a label to the pod * add a trace inject test * add downwardAPi for controlplaneTracing * add controlPlaneTracing condition to volumeMounts * update add-ons to have workload-ns * add workload-ns label to control-plane components Signed-off-by: Tarun Pothulapati <tarunpothulapati@outlook.com>	2020-04-22 10:33:51 -05:00
Alex Leong	9bf54d36ed	Upgrade to go 1.14.2 (#4278 ) Upgrade Linkerd's base docker image to use go 1.14.2 in order to stay modern. The only code change required was to update a test which was checking the error message of a `crypto/x509.CertificateInvalidError`. The error message of this error changed between go versions. We update the test to not check for the specific error string so that this test passes regardless of go version. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-04-20 17:14:51 -07:00
Alex Leong	5d3862c120	Use /live for liveness probe (#4270 ) Fixes #3984 We use the new `/live` admin endpoint in the Linkerd proxy for liveness probes instead of the `/metrics` endpoint. This endpoint returns a much smaller payload. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-04-17 14:53:32 -07:00
Kevin Leimkuhler	b6aad75b35	Add `operationID` field to tap openapi response (#4245 ) This fixes an issue users are experiencing when upgrading from from Linkerd 2.6 to 2.7 and use the [kubernetes-external-secrets]() project. The change introduced by #3700 resulted in the tap service showing up in the `/openapi/v2` API response. I confirmed this with a local build. A dependency within the project expects the `operationID` field to be present in the swagger definition. It is optional as stated in the [spec](https://swagger.io/docs/specification/paths-and-operations/). It's purpose is to identify an operation and should be unique. This change adds that field to tap service swagger spec. While this can be fixed in the KES dependency, it certainly does not hurt to add and other libraries may similarly expect this field. Signed-off-by: Kevin Leimkuhler <kevin@kleimkuhler.com>	2020-04-15 09:41:06 -07:00
Zahari Dichev	26c14d3c66	Detect changes in addresses when getting updates in endpoints watcher (#4104 ) Detect changes in addresses when getting updates in endpoints watcher Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-04-10 11:42:39 +03:00
Alex Leong	d8eebee4f7	Upgrade to client-go 0.17.4 and smi-sdk-go 0.3.0 (#4221 ) Here we upgrade our dependencies on client-go to 0.17.4 and smi-sdk-go to 0.3.0. Since smi-sdk-go uses client-go 0.17.4, these upgrades must be performed simultaneously. This also requires simultaneously upgrading our dependency on linkerd/stern to a SHA which also uses client-go 0.17.4. This keeps all of our transitive dependencies synchronized on one version of client-go. This ALSO requires updating our codegen scripts to use the 0.17.4 version of code-generator and running it to generate 0.17.4 compatible generated code. I took this opportunity to update our code generation script to properly use the version of code-generater from `go.mod` rather than a hardcoded SHA. Signed-off-by: Alex Leong <alex@buoyant.io>	2020-04-01 10:07:23 -07:00
Zahari Dichev	10ecd8889e	Set auth override (#4160 ) Set AuthOverride when present on endpoints annotation Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-03-25 10:56:36 +02:00
Mayank Shah	963b9b049a	Add kubectl-style label selectors (#4120 ) * Update tap, routes and top commands to support label selectors Signed-off-by: Mayank Shah <mayankshah1614@gmail.com>	2020-03-20 10:45:06 -05:00
Alejandro Pedraza	8f79e07ee2	Bump proxy-init to v1.3.2 (#4170 ) * Bump proxy-init to v1.3.2 Bumped `proxy-init` version to v1.3.2, fixing an issue with `go.mod` (linkerd/linkerd2-proxy-init#9). This is a non-user-facing fix.	2020-03-17 14:49:25 -05:00
Kevin Leimkuhler	10db65bcb3	Update linkerd/stern to fix go.mod parsing (#4173 ) ## Motivation I noticed the Go language server stopped working in VS Code and narrowed it down to `go build ./...` failing with the following: ``` ❯ go build ./... go: github.com/linkerd/stern@v0.0.0-20190907020106-201e8ccdff9c: parsing go.mod: go.mod:3: usage: go 1.23 ``` This change updates `linkerd/stern` version with changes made in linkerd/stern#3 to fix this issue. This does not depend on #4170, but it is also needed in order to completely fix `go build ./...`	2020-03-17 11:16:18 -07:00
Zahari Dichev	2db307ee91	Remove target port requirement in port resolution (#4174 ) This change removes the target port requirement when resolving ports in the dst service. Based on the comments, it seems that we need to have a target port defined in the port spec in order to resolve to the port in the Endpoints. In reality if target port is note defined when creating the service, k8s will set the port and the target port to the same value. Seems to me that checking for the targetPort to be different than 0, is a no-op. Signed-off-by: Zahari Dichev zaharidichev@gmail.com	2020-03-16 23:04:08 +02:00
Zahari Dichev	caf4e61daf	Enable identitiy on endpoints not associated with pods (#4134 ) Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-03-09 20:55:57 +02:00
Zahari Dichev	72fc94b03c	Service mirroring tests (#4115 ) Unit tests that exercise most of the code in cluster_watcher.go. Essentially the whole cluster mirroring machinary can be tought of as a function that takes remote cluster state, local cluster state, and modification events and as a result it either modifies local cluster state or issues new events onto the queue. This is what these tests are trying to model. I think this covers a lot of the logic there. Any suggestions for other edge cases are welcome. Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-03-04 20:17:21 +02:00
Zahari Dichev	edd7fd203d	Service Mirroring Component (#4028 ) This PR introduces a service mirroring component that is responsible for watching remote clusters and mirroring their services locally. Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-03-02 21:16:08 +02:00
Christy Jacob	8111e54606	Check for extension server certificate (#4062 ) * Check Extension api server Authentication * Added Checks and tests for extension api-server authentication * Fixed Failing Static Checks * Updated the golden file Signed-off-by: Christy Jacob <christyjacob4@gmail.com>	2020-02-28 13:39:02 -08:00
Mayank Shah	3c3a4a5f5d	cli: Add label selector flag for `stat` (#4040 ) * Update `linkerd-namespace` shorthand to `L` * Add --selector (-l) flag for `stat` Signed-off-by: Mayank Shah <mayankshah1614@gmail.com>	2020-02-17 13:40:07 -05:00
Zahari Dichev	6fa9407318	Ensure we get the correct type out of Informer Deletion events (#4034 ) Ensure we get what we expect when receiving DELETE events from the k8s Informer api Signed-off-by: Zahari Dichev <zaharidichev@gmail.com>	2020-02-15 10:15:24 +02:00
Alex Leong	ec51434eb9	Show traffic split metrics from sources in all namespaces (#3967 ) Fixes #3562 When a pod in one namespace sends traffic to a service which is the apex of a traffic split in another namespace, that traffic is not displayed in the `linkerd stat trafficsplit` output. This is because when we do a Prometheus query for traffic to the traffic split, we supply a Prometheus label selector to only select traffic sources in the namespace of the traffic split. Since any pod in any namespace can send traffic to the apex service of a traffic split, we must look at all possible sources of traffic, not just the ones in the same namespace. Before: ``` $ bin/linkerd stat ts NAME APEX LEAF WEIGHT SUCCESS RPS LATENCY_P50 LATENCY_P95 LATENCY_P99 webapp-split webapp webapp 900m - - - - - webapp-split webapp webapp-2 100m - - - - - ``` After: ``` $ bin/linkerd stat ts NAME APEX LEAF WEIGHT SUCCESS RPS LATENCY_P50 LATENCY_P95 LATENCY_P99 webapp-split webapp webapp 900m 80.00% 1.4rps 31ms 99ms 2530ms webapp-split webapp webapp-2 100m 60.00% 0.2rps 35ms 93ms 99ms ``` Signed-off-by: Alex Leong <alex@buoyant.io>	2020-02-12 09:21:59 -08:00
Alejandro Pedraza	3ba66f6f9d	Fix flakey TestGetProfiles (#3965 ) Fixes #3332 Fixes the very rare test failure ``` --- FAIL: TestGetProfiles (0.33s) --- FAIL: TestGetProfiles/Returns_server_profile (0.11s) server_test.go:228: Expected 1 or 2 updates but got 3: [retry_budget:<retry_ratio:0.2 min_retries_per_second:10 ttl:<seconds:10 > > routes:<condition:<path:<regex:"/a/b/c" > > metrics_labels:<key:"route" value:"route1" > timeout:<seconds:10 > > retry_budget:<retry_ratio:0.2 min_retries_per_second:10 ttl:<seconds:10 > > routes:<condition:<path:<regex:"/a/b/c" > > metrics_labels:<key:"route" value:"route1" > timeout:<seconds:10 > > retry_budget:<retry_ratio:0.2 min_retries_per_second:10 ttl:<seconds:10 > > ] FAIL FAIL github.com/linkerd/linkerd2/controller/api/destination 0.624s ``` that occurs when a third unexpected stream update occurs, when the fake API takes more time to notify its listeners about the resources created. For all the nasty details check #3332	2020-02-07 19:43:29 -05:00
Dax McDonald	76d3285247	Use correct go module file syntax (#4021 ) The correct syntax for the go module file is go MAJOR.MINOR Signed-off-by: Dax McDonald <dax@rancher.com>	2020-02-07 07:58:54 -08:00
Alejandro Pedraza	afb93cddc8	Use `t.Name()` instead of `t.Name` in tests (#3970 ) Use `t.Name()` instead of `t.Name` when retrieving the name of tests. This was causing an error to be added in the log: ``` output: logrus_error="can not add field \"test\" ``` Followup to [comment](https://github.com/linkerd/linkerd2/pull/3965#discussion_r370387990)	2020-01-27 09:17:19 -05:00

1 2 3 4 5 ...

498 Commits